Voice interaction services
Abstract
The disclosed embodiments include computerized methods, systems, and devices, including computer programs encoded on a computer storage medium, for integrating voice-based interaction and control into a native graphical user interface (GUI) of an executed application. For example, a communications device may receive audio data corresponding to an utterance spoken by a user, and may obtain structured data representative of the received audio data. The communications device may provide structured data to the executed application through a programmatic interface, and the executed application may perform the one or more operations in accordance with the structured data. The communications device may generate data indicative of an output of the one or more operations performed by the executed application, and may present at least a portion of the generated output data to a user through a corresponding interface.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method, comprising:
obtaining, at a communications device, and by one or more processors, structured data representative of utterance spoken by a user, the utterance specifying a functionality of an application executed by the communications device, and the structured data causing the executed application to perform one or more operations consistent with the specified functionality; providing, by the one or more processors, the structured data to the executed application through a programmatic interface, the executed application performing the one or more operations in accordance with the structured data; generating, by the one or more processors, data indicative of an output of the one or more operations performed by the executed application; and presenting, by the one or more processors, at least a portion of the generated output data to a user through a corresponding interface.
2 . The method of claim 1 , wherein the utterance is spoken by the user into a microphone of an additional communications device.
3 . The method of claim 1 , further comprising receiving audio data corresponding to the utterance at the communications device, the utterance being spoken by the user into a microphone of the communications device, and the obtained structured data being representative of the received audio data.
4 . The method of claim 3 , wherein:
the executed application corresponds to a foreground application; and the obtaining further comprises executing a background application to generate the portion of the structured data.
5 . The method of claim 3 , further comprising:
presenting, to the user through the corresponding interface, an interface element identifying the microphone; receiving input from the user indicative of a selection of the presented interface element; and performing operations that activate the microphone in response to the received input, the activated microphone being configured to capture the utterance spoken by the user; and modifying a visual characteristic of the presented interface element in response to the received input, the modified visual characteristic being perceptible by the user and being indicative of the activation of the microphone.
6 . The method of claim 3 , further comprising:
transmitting a portion of the audio data to an external computing system, the external computing system being configured to generate the structured data based on an application of one or more of a speech recognition algorithm and a natural-language processing algorithm to the transmitted portion of the audio data; and receiving the structured data from the external computing system across the corresponding communications network.
7 . The method of claim 6 , further comprising:
obtaining contextual data indicative of an interaction of the user with the executed application; and transmitting portions of the audio data and the contextual data to the external computing system, the external computing system being configured to generate the structured data based on an application of one or more of a speech recognition algorithm and a natural-language processing algorithm to the transmitted portions of the audio data and the contextual data.
8 . The method of claim 3 , wherein the method further comprises:
applying one or more of a speech recognition algorithm and a natural-language processing algorithm to the received audio data; and generating at least a portion of the structured data representative of the received audio data based on the application of the one or more of the speech recognition algorithm or the natural-language processing algorithm.
9 . The method of claim 3 , wherein the method further comprises:
applying one or more of a semantic parsing algorithm and speech biasing technique to the received audio data; and generating at least one a portion of the structured data representative of the received audio data based on the application of the one or more of the speech recognition algorithm or the natural-language processing algorithm.
10 . The method of claim 3 , further comprising:
obtaining contextual data indicative of an interaction of the user with the executed application; and applying the speech recognition algorithm to portions of the received audio data and the obtained contextual data; based on the application of the speech recognition algorithm, identifying linguistic elements that correspond to the utterance spoken by the user; applying the natural language processing algorithm to the identified linguistic elements and to data identifying a plurality of actions associated with the executed application; based on the application of the at least one natural language algorithm, determining that the identified linguistic elements correspond to at least one of the actions associated with the executed application; determining a format associated with the at least one of the actions, the determined format identifying one or more data inputs associated with the at least one action; and generating at least a portion of the structured data representative of the received audio data in accordance with the determined format, the generated portion of the structured data identifying the at least one action and the data inputs.
11 . The method of claim 1 , wherein:
the structured data corresponds to a command associated with the executed application; and the structured data comprises one or more data fields, the data fields identifying at least one of a command type, a command sub-type, or a characteristic of an element of digital content associated with the command.
12 . The method of claim 1 , wherein:
the communications device further comprises a speaker; and the method further comprises:
presenting, through the speaker, audible content associated with at least one of the utterance or the executed application; and
in response to the presented audible content, receiving additional audio data at the communications device,
the audio data corresponding to an additional utterance spoken by the user into the microphone.
13 . A communications device, comprising:
at least one processor; and a memory storing executable instructions that, when executed by the at least one processor, causes the at least one processor to perform the steps of:
obtaining structured data representative of an utterance spoken by a user, the utterance specifying a functionality of an application executed by the communications device, and the structured data causing the executed application to perform one or more operations consistent with the specified functionality;
providing the structured data to the executed application through a programmatic interface, the executed application performing the one or more operations in accordance with the structured data;
generating data indicative of an output of the one or more operations performed by the executed application; and
presenting at least a portion of the generated output data to a user through a corresponding interface.
14 . The communications device of claim 13 , wherein the utterance is spoken by the user into a microphone of an additional communications device.
15 . The communications device of claim 13 , wherein:
the communications device further comprises a microphone coupled to the at least one processor; and the at least one processor further performs the step of receiving audio data corresponding to the utterance at the communications device, the utterance being spoken by the user into the microphone, and the obtained structured data being representative of the received audio data.
16 . The communications device of claim 15 , wherein the at least one processor further performs the steps of:
presenting, to the user through the corresponding interface, an interface element identifying the microphone; receiving input from the user indicative of a selection of the presented interface element; performing operations that activate the microphone in response to the received input, the activated microphone being configured to capture the utterance spoken by the user; and modifying a visual characteristic of the presented interface element in response to the received input, the modified visual characteristic being perceptible by the user and being indicative of the activation of the microphone.
17 . The communications device of claim 15 , wherein the at least one processor further performs the steps of:
transmitting a portion of the audio data to an external computing system, the external computing system being configured to generate the structured data based on an application of one or more of a speech recognition algorithm and a natural-language processing algorithm to the transmitted portion of the audio data; and receiving the structured data from the external computing system across the corresponding communications network.
18 . The communications device of claim 17 , wherein the at least one processor further performs the steps of:
obtaining contextual data indicative of an interaction of the user with the executed application; and transmitting portions of the audio data and the contextual data to the external computing system, the external computing system being configured to generate the structured data based on an application of one or more of a speech recognition algorithm and a natural-language processing algorithm to the transmitted portions of the audio data and the contextual data.
19 . The communications device of claim 15 , wherein:
the at least one processor further performs the step of generating at least a portion of the structured data representative of the received audio data; and the step of generating the portion of the structured data comprises:
applying one or more of a speech recognition algorithm, a natural-language processing algorithm, a semantic parsing algorithm, and a speech biasing technique to the received audio data; and
generating the structured representation based on the application of the one or more of the speech recognition algorithm or the natural-language processing algorithm.
20 . A system, comprising:
one or more computers; and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
obtaining audio data captured by a microphone of a communications device, the audio data corresponding to an utterance spoken by a user into the microphone, and the utterance specifying a functionality of an executable application;
applying one or more of a speech recognition algorithm and a natural-language processing algorithm to the obtained audio data;
based on the application of the one or more of the speech recognition algorithm or the natural-language processing algorithm, generating structured data representative of the received audio data; and
transmitting at least a portion of the structured data to a destination device, the destination device executing the executable application, and the structured data causing the executable application to perform one or more operations consistent with the specified functionality.
21 . The system of claim 20 , wherein:
the destination device corresponds to the communications device; and the at least one processor further performs the step of transmitting at least the portion of the structured data to the communications device, the communications device executing the executable application.
22 . The system of claim 20 , wherein the at least one processor further performs the steps of:
accessing data identifying a plurality of candidate devices, the candidate devices comprising at least one of an additional communications device, a wearable device, a connected device, or a vehicle-based device, and the accessed data associating the candidate devices with corresponding characteristics of the received audio data; and based on portions of the received audio data and the accessed data, establishing a corresponding one of the candidate devices as the destination device, the corresponding one of the candidate destination devices being configured to execute the executable application.
23 . The system of claim 20 , wherein the one or more computers further perform the operations of:
obtaining contextual data indicative of an interaction of the user with the executed application; applying the speech recognition algorithm to portions of the received audio data and the obtained contextual data; based on the application of the speech recognition algorithm, identifying linguistic elements that correspond to the utterance spoken by the user; applying the natural language processing algorithm to the identified linguistic elements and to data identifying a plurality of actions associated with the executed application; based on the application of the at least one natural language algorithm, determining that the identified linguistic elements correspond to at least one of the actions associated with the executed application; determining a format associated with the at least one of the actions, the determined format identifying one or more data inputs associated with the at least one action; and generating the portion of the structured data accordance with the determined format, the generated portion of the structured data identifying the at least one action and the data inputs.Join the waitlist — get patent alerts
Track US2018039478A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.