US2018039478A1PendingUtilityA1

Voice interaction services

Assignee: GOOGLE INCPriority: Aug 2, 2016Filed: Aug 2, 2016Published: Feb 8, 2018
Est. expiryAug 2, 2036(~10 yrs left)· nominal 20-yr term from priority
G10L 2015/223G10L 15/30G06F 3/167G06F 3/04842G10L 2015/228G06F 3/0481G10L 15/1822G10L 2015/225G10L 15/22G10L 15/1815
43
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The disclosed embodiments include computerized methods, systems, and devices, including computer programs encoded on a computer storage medium, for integrating voice-based interaction and control into a native graphical user interface (GUI) of an executed application. For example, a communications device may receive audio data corresponding to an utterance spoken by a user, and may obtain structured data representative of the received audio data. The communications device may provide structured data to the executed application through a programmatic interface, and the executed application may perform the one or more operations in accordance with the structured data. The communications device may generate data indicative of an output of the one or more operations performed by the executed application, and may present at least a portion of the generated output data to a user through a corresponding interface.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method, comprising:
 obtaining, at a communications device, and by one or more processors, structured data representative of utterance spoken by a user, the utterance specifying a functionality of an application executed by the communications device, and the structured data causing the executed application to perform one or more operations consistent with the specified functionality;   providing, by the one or more processors, the structured data to the executed application through a programmatic interface, the executed application performing the one or more operations in accordance with the structured data;   generating, by the one or more processors, data indicative of an output of the one or more operations performed by the executed application; and   presenting, by the one or more processors, at least a portion of the generated output data to a user through a corresponding interface.   
     
     
         2 . The method of  claim 1 , wherein the utterance is spoken by the user into a microphone of an additional communications device. 
     
     
         3 . The method of  claim 1 , further comprising receiving audio data corresponding to the utterance at the communications device, the utterance being spoken by the user into a microphone of the communications device, and the obtained structured data being representative of the received audio data. 
     
     
         4 . The method of  claim 3 , wherein:
 the executed application corresponds to a foreground application; and   the obtaining further comprises executing a background application to generate the portion of the structured data.   
     
     
         5 . The method of  claim 3 , further comprising:
 presenting, to the user through the corresponding interface, an interface element identifying the microphone;   receiving input from the user indicative of a selection of the presented interface element; and   performing operations that activate the microphone in response to the received input, the activated microphone being configured to capture the utterance spoken by the user; and   modifying a visual characteristic of the presented interface element in response to the received input, the modified visual characteristic being perceptible by the user and being indicative of the activation of the microphone.   
     
     
         6 . The method of  claim 3 , further comprising:
 transmitting a portion of the audio data to an external computing system, the external computing system being configured to generate the structured data based on an application of one or more of a speech recognition algorithm and a natural-language processing algorithm to the transmitted portion of the audio data; and   receiving the structured data from the external computing system across the corresponding communications network.   
     
     
         7 . The method of  claim 6 , further comprising:
 obtaining contextual data indicative of an interaction of the user with the executed application; and   transmitting portions of the audio data and the contextual data to the external computing system, the external computing system being configured to generate the structured data based on an application of one or more of a speech recognition algorithm and a natural-language processing algorithm to the transmitted portions of the audio data and the contextual data.   
     
     
         8 . The method of  claim 3 , wherein the method further comprises:
 applying one or more of a speech recognition algorithm and a natural-language processing algorithm to the received audio data; and   generating at least a portion of the structured data representative of the received audio data based on the application of the one or more of the speech recognition algorithm or the natural-language processing algorithm.   
     
     
         9 . The method of  claim 3 , wherein the method further comprises:
 applying one or more of a semantic parsing algorithm and speech biasing technique to the received audio data; and   generating at least one a portion of the structured data representative of the received audio data based on the application of the one or more of the speech recognition algorithm or the natural-language processing algorithm.   
     
     
         10 . The method of  claim 3 , further comprising:
 obtaining contextual data indicative of an interaction of the user with the executed application; and   applying the speech recognition algorithm to portions of the received audio data and the obtained contextual data;   based on the application of the speech recognition algorithm, identifying linguistic elements that correspond to the utterance spoken by the user;   applying the natural language processing algorithm to the identified linguistic elements and to data identifying a plurality of actions associated with the executed application;   based on the application of the at least one natural language algorithm, determining that the identified linguistic elements correspond to at least one of the actions associated with the executed application;   determining a format associated with the at least one of the actions, the determined format identifying one or more data inputs associated with the at least one action; and   generating at least a portion of the structured data representative of the received audio data in accordance with the determined format, the generated portion of the structured data identifying the at least one action and the data inputs.   
     
     
         11 . The method of  claim 1 , wherein:
 the structured data corresponds to a command associated with the executed application; and   the structured data comprises one or more data fields, the data fields identifying at least one of a command type, a command sub-type, or a characteristic of an element of digital content associated with the command.   
     
     
         12 . The method of  claim 1 , wherein:
 the communications device further comprises a speaker; and   the method further comprises:
 presenting, through the speaker, audible content associated with at least one of the utterance or the executed application; and 
 in response to the presented audible content, receiving additional audio data at the communications device, 
   the audio data corresponding to an additional utterance spoken by the user into the microphone.   
     
     
         13 . A communications device, comprising:
 at least one processor; and   a memory storing executable instructions that, when executed by the at least one processor, causes the at least one processor to perform the steps of:
 obtaining structured data representative of an utterance spoken by a user, the utterance specifying a functionality of an application executed by the communications device, and the structured data causing the executed application to perform one or more operations consistent with the specified functionality; 
 providing the structured data to the executed application through a programmatic interface, the executed application performing the one or more operations in accordance with the structured data; 
 generating data indicative of an output of the one or more operations performed by the executed application; and 
 presenting at least a portion of the generated output data to a user through a corresponding interface. 
   
     
     
         14 . The communications device of  claim 13 , wherein the utterance is spoken by the user into a microphone of an additional communications device. 
     
     
         15 . The communications device of  claim 13 , wherein:
 the communications device further comprises a microphone coupled to the at least one processor; and   the at least one processor further performs the step of receiving audio data corresponding to the utterance at the communications device, the utterance being spoken by the user into the microphone, and the obtained structured data being representative of the received audio data.   
     
     
         16 . The communications device of  claim 15 , wherein the at least one processor further performs the steps of:
 presenting, to the user through the corresponding interface, an interface element identifying the microphone;   receiving input from the user indicative of a selection of the presented interface element;   performing operations that activate the microphone in response to the received input, the activated microphone being configured to capture the utterance spoken by the user; and   modifying a visual characteristic of the presented interface element in response to the received input, the modified visual characteristic being perceptible by the user and being indicative of the activation of the microphone.   
     
     
         17 . The communications device of  claim 15 , wherein the at least one processor further performs the steps of:
 transmitting a portion of the audio data to an external computing system, the external computing system being configured to generate the structured data based on an application of one or more of a speech recognition algorithm and a natural-language processing algorithm to the transmitted portion of the audio data; and   receiving the structured data from the external computing system across the corresponding communications network.   
     
     
         18 . The communications device of  claim 17 , wherein the at least one processor further performs the steps of:
 obtaining contextual data indicative of an interaction of the user with the executed application; and   transmitting portions of the audio data and the contextual data to the external computing system, the external computing system being configured to generate the structured data based on an application of one or more of a speech recognition algorithm and a natural-language processing algorithm to the transmitted portions of the audio data and the contextual data.   
     
     
         19 . The communications device of  claim 15 , wherein:
 the at least one processor further performs the step of generating at least a portion of the structured data representative of the received audio data; and   the step of generating the portion of the structured data comprises:
 applying one or more of a speech recognition algorithm, a natural-language processing algorithm, a semantic parsing algorithm, and a speech biasing technique to the received audio data; and 
 generating the structured representation based on the application of the one or more of the speech recognition algorithm or the natural-language processing algorithm. 
   
     
     
         20 . A system, comprising:
 one or more computers; and   one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
 obtaining audio data captured by a microphone of a communications device, the audio data corresponding to an utterance spoken by a user into the microphone, and the utterance specifying a functionality of an executable application; 
 applying one or more of a speech recognition algorithm and a natural-language processing algorithm to the obtained audio data; 
 based on the application of the one or more of the speech recognition algorithm or the natural-language processing algorithm, generating structured data representative of the received audio data; and 
 transmitting at least a portion of the structured data to a destination device, the destination device executing the executable application, and the structured data causing the executable application to perform one or more operations consistent with the specified functionality. 
   
     
     
         21 . The system of  claim 20 , wherein:
 the destination device corresponds to the communications device; and   the at least one processor further performs the step of transmitting at least the portion of the structured data to the communications device, the communications device executing the executable application.   
     
     
         22 . The system of  claim 20 , wherein the at least one processor further performs the steps of:
 accessing data identifying a plurality of candidate devices, the candidate devices comprising at least one of an additional communications device, a wearable device, a connected device, or a vehicle-based device, and the accessed data associating the candidate devices with corresponding characteristics of the received audio data; and   based on portions of the received audio data and the accessed data, establishing a corresponding one of the candidate devices as the destination device, the corresponding one of the candidate destination devices being configured to execute the executable application.   
     
     
         23 . The system of  claim 20 , wherein the one or more computers further perform the operations of:
 obtaining contextual data indicative of an interaction of the user with the executed application;   applying the speech recognition algorithm to portions of the received audio data and the obtained contextual data;   based on the application of the speech recognition algorithm, identifying linguistic elements that correspond to the utterance spoken by the user;   applying the natural language processing algorithm to the identified linguistic elements and to data identifying a plurality of actions associated with the executed application;   based on the application of the at least one natural language algorithm, determining that the identified linguistic elements correspond to at least one of the actions associated with the executed application;   determining a format associated with the at least one of the actions, the determined format identifying one or more data inputs associated with the at least one action; and   generating the portion of the structured data accordance with the determined format, the generated portion of the structured data identifying the at least one action and the data inputs.

Join the waitlist — get patent alerts

Track US2018039478A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.