US2023410816A1PendingUtilityA1

Dialog management with multiple modalities

Assignee: AMAZON TECH INCPriority: Nov 18, 2013Filed: Jun 26, 2023Published: Dec 21, 2023
Est. expiryNov 18, 2033(~7.3 yrs left)· nominal 20-yr term from priority
G10L 17/00G10L 15/22G10L 15/183G10L 15/18G10L 2015/228G10L 2015/223
73
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Features are disclosed for performing functions in response to user requests based on contextual data regarding prior user requests. Users may engage in conversations with a computing device in order to initiate some function or obtain some information. A dialog manager may manage the conversations and store contextual data regarding one or more of the conversations. Processing and responding to subsequent conversations may benefit from the previously stored contextual data by, e.g., reducing the amount of information that a user must provide if the user has already provided the information in the context of a prior conversation. Additional information associated with performing functions responsive to user requests may be shared among applications, further improving efficiency and enhancing the user experience.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A system comprising:
 computer-readable memory storing executable instructions; and   one or more processors in communication with the computer-readable memory, wherein the one or more processors are programmed by the executable instructions to:
 receive audio data representing at least one utterance; 
 generate, using natural language understanding (“NLU”) processing based at least partly on the audio data, command data that represents the at least one utterance; 
 send the command data to a first application resulting in first output content being presented in a first modality during a first period of time in response to receiving the command data; 
 receive second output content in a second modality different from the first modality, wherein the second output content is associated with the first output content and a second application; and 
 cause presentation of the second output content in the second modality during a second period of time beginning subsequent to a beginning of the first period of time. 
   
     
     
         3 . The system of  claim 2 , wherein the first modality comprises a same modality as the audio data representing the at least one utterance. 
     
     
         4 . The system of  claim 2 , wherein first output content comprises an audio presentation, and wherein the second output content comprises a visual presentation. 
     
     
         5 . The system of  claim 2 , wherein the one or more processors are programmed by further executable instructions to manage a multi-turn dialog comprising the at least one utterance, at least a second utterance, and at least one system-generated response. 
     
     
         6 . The system of  claim 2 , wherein the one or more processors are programmed by further executable instructions to determine to send the command data to the first application based at least partly on a subject of the at least one utterance. 
     
     
         7 . The system of  claim 2 , wherein the one or more processors are programmed by further executable instructions to:
 generate, using automatic speech recognition (“ASR”) processing and the audio data, utterance data representing the at least one utterance,   wherein the command data being generated using NLU processing based at least partly on the audio data comprises the command data being generated using NLU processing and the utterance data.   
     
     
         8 . The system of  claim 2 , wherein the one or more processors are programmed by further executable instructions to:
 store context data, wherein the context data is generated by the first application;   generate, using second NLU processing based at least partly on second audio data, second command data;   determine that a second application is to use the context data; and   send the second command data and the context data to the second application.   
     
     
         9 . The system of  claim 2 , wherein the one or more processors are programmed by further executable instructions to:
 generate context data during the NLU processing, wherein the context data is based on at least one of: named entity data associated with the at least one utterance, or utterance data representing a plurality of previously processed utterances;   store the context data; and   generate second command data using second NLU processing based at least partly on second audio data and the context data.   
     
     
         10 . The system of  claim 2 , further comprising:
 an audio input device configured to generate the audio data based on the at least one utterance; and   an audio output device configured to present at least one of the first output content or the second output content.   
     
     
         11 . A computer-implemented method comprising:
 under control of a computing system comprising one or more processors configured with specific computer-executable instructions,
 receiving audio data representing at least one utterance; 
 generating, using natural language understanding (“NLU”) processing based at least partly on the audio data, command data that represents the at least one utterance; 
 sending the command data to a first application resulting in first output content being presented in a first modality during a first period of time in response to receiving the command data; 
 receiving second output content in a second modality different from the first modality, wherein the second output content is associated with the first output content and a second application; and 
 causing presentation of the second output content in the second modality during a second period of time beginning subsequent to a beginning of the first period of time. 
   
     
     
         12 . The computer-implemented method of  claim 11 , further comprising obtaining, by the first application, the first output content in the first modality, wherein the first modality is a same modality as the audio data representing the at least one utterance. 
     
     
         13 . The computer-implemented method of  claim 11 , further comprising obtaining, by the second application, the second output content in the second modality, wherein the first modality is different from a modality of the audio data representing the at least one utterance. 
     
     
         14 . The computer-implemented method of  claim 11 , wherein the causing the presentation of the second output content in the second modality comprises causing presentation of the second output content in a visual mode of presentation, wherein the first modality comprises an audio mode of presentation. 
     
     
         15 . The computer-implemented method of  claim 11 , wherein sending the command data to the first application is performed by a dialog manager application, and wherein causing presentation of the second output content in the second modality comprises sending, by the dialog manager application, the second output content to the second application configured to present content in the second modality. 
     
     
         16 . The computer-implemented method of  claim 11 , wherein causing presentation of the second output content in the second modality comprises causing presentation of the second output content in the second modality using a different output device than an output device used by the first application to cause presentation of the first output content in the first modality. 
     
     
         17 . The computer-implemented method of  claim 11 , further comprising storing context data associated with the at least one utterance, wherein the context data comprises the second output content. 
     
     
         18 . The computer-implemented method of  claim 17 , further comprising:
 receiving second audio data representing a second utterance;   determining that the second utterance is part of a conversation including the at least one utterance; and   determining, based at least partly on the context data, to cause presentation of the second output content in the second modality in response to receiving the second audio data.   
     
     
         19 . The computer-implemented method of  claim 11 , further comprising determining to send the command data to the first application based at least partly on a subject of the at least one utterance. 
     
     
         20 . The computer-implemented method of  claim 11 , further comprising:
 generating context data during the NLU processing;   storing the context data; and   generating second command data using second NLU processing based at least partly on second audio data and the context data.   
     
     
         21 . The computer-implemented method of  claim 11 , where the receiving the audio data comprises receiving the audio data over a network connection to a computing device.

Join the waitlist — get patent alerts

Track US2023410816A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.