Dialog management with multiple modalities
Abstract
Features are disclosed for performing functions in response to user requests based on contextual data regarding prior user requests. Users may engage in conversations with a computing device in order to initiate some function or obtain some information. A dialog manager may manage the conversations and store contextual data regarding one or more of the conversations. Processing and responding to subsequent conversations may benefit from the previously stored contextual data by, e.g., reducing the amount of information that a user must provide if the user has already provided the information in the context of a prior conversation. Additional information associated with performing functions responsive to user requests may be shared among applications, further improving efficiency and enhancing the user experience.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A system comprising:
computer-readable memory storing executable instructions; and one or more processors in communication with the computer-readable memory, wherein the one or more processors are programmed by the executable instructions to:
receive audio data representing at least one utterance;
generate, using natural language understanding (“NLU”) processing based at least partly on the audio data, command data that represents the at least one utterance;
send the command data to a first application resulting in first output content being presented in a first modality during a first period of time in response to receiving the command data;
receive second output content in a second modality different from the first modality, wherein the second output content is associated with the first output content and a second application; and
cause presentation of the second output content in the second modality during a second period of time beginning subsequent to a beginning of the first period of time.
3 . The system of claim 2 , wherein the first modality comprises a same modality as the audio data representing the at least one utterance.
4 . The system of claim 2 , wherein first output content comprises an audio presentation, and wherein the second output content comprises a visual presentation.
5 . The system of claim 2 , wherein the one or more processors are programmed by further executable instructions to manage a multi-turn dialog comprising the at least one utterance, at least a second utterance, and at least one system-generated response.
6 . The system of claim 2 , wherein the one or more processors are programmed by further executable instructions to determine to send the command data to the first application based at least partly on a subject of the at least one utterance.
7 . The system of claim 2 , wherein the one or more processors are programmed by further executable instructions to:
generate, using automatic speech recognition (“ASR”) processing and the audio data, utterance data representing the at least one utterance, wherein the command data being generated using NLU processing based at least partly on the audio data comprises the command data being generated using NLU processing and the utterance data.
8 . The system of claim 2 , wherein the one or more processors are programmed by further executable instructions to:
store context data, wherein the context data is generated by the first application; generate, using second NLU processing based at least partly on second audio data, second command data; determine that a second application is to use the context data; and send the second command data and the context data to the second application.
9 . The system of claim 2 , wherein the one or more processors are programmed by further executable instructions to:
generate context data during the NLU processing, wherein the context data is based on at least one of: named entity data associated with the at least one utterance, or utterance data representing a plurality of previously processed utterances; store the context data; and generate second command data using second NLU processing based at least partly on second audio data and the context data.
10 . The system of claim 2 , further comprising:
an audio input device configured to generate the audio data based on the at least one utterance; and an audio output device configured to present at least one of the first output content or the second output content.
11 . A computer-implemented method comprising:
under control of a computing system comprising one or more processors configured with specific computer-executable instructions,
receiving audio data representing at least one utterance;
generating, using natural language understanding (“NLU”) processing based at least partly on the audio data, command data that represents the at least one utterance;
sending the command data to a first application resulting in first output content being presented in a first modality during a first period of time in response to receiving the command data;
receiving second output content in a second modality different from the first modality, wherein the second output content is associated with the first output content and a second application; and
causing presentation of the second output content in the second modality during a second period of time beginning subsequent to a beginning of the first period of time.
12 . The computer-implemented method of claim 11 , further comprising obtaining, by the first application, the first output content in the first modality, wherein the first modality is a same modality as the audio data representing the at least one utterance.
13 . The computer-implemented method of claim 11 , further comprising obtaining, by the second application, the second output content in the second modality, wherein the first modality is different from a modality of the audio data representing the at least one utterance.
14 . The computer-implemented method of claim 11 , wherein the causing the presentation of the second output content in the second modality comprises causing presentation of the second output content in a visual mode of presentation, wherein the first modality comprises an audio mode of presentation.
15 . The computer-implemented method of claim 11 , wherein sending the command data to the first application is performed by a dialog manager application, and wherein causing presentation of the second output content in the second modality comprises sending, by the dialog manager application, the second output content to the second application configured to present content in the second modality.
16 . The computer-implemented method of claim 11 , wherein causing presentation of the second output content in the second modality comprises causing presentation of the second output content in the second modality using a different output device than an output device used by the first application to cause presentation of the first output content in the first modality.
17 . The computer-implemented method of claim 11 , further comprising storing context data associated with the at least one utterance, wherein the context data comprises the second output content.
18 . The computer-implemented method of claim 17 , further comprising:
receiving second audio data representing a second utterance; determining that the second utterance is part of a conversation including the at least one utterance; and determining, based at least partly on the context data, to cause presentation of the second output content in the second modality in response to receiving the second audio data.
19 . The computer-implemented method of claim 11 , further comprising determining to send the command data to the first application based at least partly on a subject of the at least one utterance.
20 . The computer-implemented method of claim 11 , further comprising:
generating context data during the NLU processing; storing the context data; and generating second command data using second NLU processing based at least partly on second audio data and the context data.
21 . The computer-implemented method of claim 11 , where the receiving the audio data comprises receiving the audio data over a network connection to a computing device.Join the waitlist — get patent alerts
Track US2023410816A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.