US2021166686A1PendingUtilityA1

Speech-based attention span for voice user interface

Assignee: AMAZON TECH INCPriority: Sep 1, 2017Filed: Oct 14, 2020Published: Jun 3, 2021
Est. expirySep 1, 2037(~11.1 yrs left)· nominal 20-yr term from priority
G10L 2015/088G10L 2015/223G10L 15/08G10L 15/22G10L 15/30G10L 25/78G10L 15/1822
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Techniques for enabling a device to send to a speech processing server further input audio data following a completed utterance dialog to prevent the need for subsequent keywords to be spoken to invoke subsequent commands are described. A system receives input audio data corresponding to an utterance from a device upon the device detecting speech corresponding to a keyword. The system performs speech processing on the input audio data to determine a command. The system determines output data responsive to the command and sends same to the device, thus completing operations regarding the utterance. The system may also send an instruction to the device to: send to the system further input audio data corresponding to further input audio without the device first detecting a wake command.

Claims

exact text as granted — not AI-modified
1 .- 20 . (canceled) 
     
     
         21 . A computer-implemented method, comprising:
 receiving first data corresponding to at least one image representing a user;   receiving audio data corresponding to an utterance spoken by the user;   processing the first data to determine that the utterance is directed at a device; and   in response to processing the first data to determine that the utterance is directed at the device, causing speech processing to be performed using the audio data.   
     
     
         22 . The computer-implemented method of  claim 21 , further comprising:
 processing the first data to determine the user is facing the device.   
     
     
         23 . The computer-implemented method of  claim 21 , further comprising:
 receiving image data representing the at least one image; and   processing the image data using a first component to determine feature data corresponding to the at least one image, wherein the first data includes the feature data,   wherein processing the first data to determine that the utterance is directed at a device comprises processing the feature data using at least one classifier.   
     
     
         24 . The computer-implemented method of  claim 21 , further comprising:
 processing the audio data to determine feature data corresponding to the utterance, wherein processing the first data to determine that the utterance is directed at a device comprises processing the first data and the feature data using at least one classifier.   
     
     
         25 . The computer-implemented method of  claim 24 , wherein processing the audio data to determine feature data comprises:
 performing automatic speech recognition (ASR) on the audio data to determine ASR result data; and   processing the ASR result data to determine the feature data.   
     
     
         26 . The computer-implemented method of  claim 21 , wherein causing speech processing to be performed using the audio data comprises sending the audio data to at least one remote device for the speech processing. 
     
     
         27 . The computer-implemented method of  claim 21 , wherein audio data was received without detection of a wakeword associated with the utterance. 
     
     
         28 . The computer-implemented method of  claim 21 , wherein audio data was received based at least in part on detection of a wakeword associated with the utterance. 
     
     
         29 . The computer-implemented method of  claim 21 , further comprising:
 processing the first data to determine the user is looking at a second device; and   based at least in part on the user looking at the second device, causing output data to be sent to the second device.   
     
     
         30 . The computer-implemented method of  claim 29 , wherein determination that the user is looking at the second device occurs after receipt of the audio data. 
     
     
         31 . A system comprising:
 at least one processor; and   at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
 receive first data corresponding to at least one image representing a user; 
 receive audio data corresponding to an utterance spoken by the user; 
 process the first data to determine that the utterance is directed at a device; and 
 in response to processing the first data to determine that the utterance is directed at the device, cause speech processing to be performed using the audio data. 
   
     
     
         32 . The system of  claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 process the first data to determine the user is facing the device.   
     
     
         33 . The system of  claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 receive image data representing the at least one image; and   process the image data using a first component to determine feature data corresponding to the at least one image, wherein the first data includes the feature data,   wherein the instructions that cause the system to process the first data to determine that the utterance is directed at a device comprise instructions that, when executed by the at least one processor, further cause the system to process the feature data using at least one classifier.   
     
     
         34 . The system of  claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 process the audio data to determine feature data corresponding to the utterance,   wherein the instructions that cause the system to process the first data to determine that the utterance is directed at a device comprise instructions that, when executed by the at least one processor, further cause the system to process the first data and the feature data using at least one classifier.   
     
     
         35 . The system of  claim 34 , the instructions that cause the system to process the audio data to determine feature data comprise instructions that, when executed by the at least one processor, further cause the system to:
 perform automatic speech recognition (ASR) on the audio data to determine ASR result data; and   process the ASR result data to determine the feature data.   
     
     
         36 . The system of  claim 31 , wherein the instructions that cause the system to cause speech processing to be performed using the audio data comprise instructions that, when executed by the at least one processor, further cause the system to send the audio data to at least one remote device for the speech processing. 
     
     
         37 . The system of  claim 31 , wherein audio data was received without detection of a wakeword associated with the utterance. 
     
     
         38 . The system of  claim 31 , wherein audio data was received based at least in part on detection of a wakeword associated with the utterance. 
     
     
         39 . The system of  claim 31 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
 process the first data to determine the user is looking at a second device; and   based at least in part on the user looking at the second device, cause output data to be sent to the second device.   
     
     
         40 . The system of  claim 39 , wherein determination that the user is looking at the second device occurs after receipt of the audio data.

Join the waitlist — get patent alerts

Track US2021166686A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.