Auto-completion for Multi-modal User Input in Assistant Systems
Abstract
In one embodiment, a method includes receiving an initial input in a first modality from a first user at a client system, determining intents and slots corresponding to the initial input, wherein the slots are conditioned on the intents, generating one or more candidate continuation-inputs based on the intents and slots, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, wherein the candidate modalities are different from the first modality, and wherein each of the candidate continuation-inputs references entities represented by the slots, and presenting one or more suggested inputs corresponding to one or more of the candidate continuation-inputs at the client system.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising, by a client system:
receiving, at the client system, an initial input from a first user, wherein the initial input is in a first modality; determining one or more intents and one or more slots corresponding to the initial input, wherein the one or more slots are conditioned on the one or more intents; generating, based on the one or more intents and the one or more slots, one or more candidate continuation-inputs, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, wherein the candidate modalities are different from the first modality, and wherein each of the candidate continuation-inputs references one or more entities represented by the one or more slots; and presenting, at the client system, one or more suggested inputs corresponding to one or more of the candidate continuation-inputs.
2 . The method of claim 1 , wherein the first modality comprises one of audio, text, image, video, motion, or orientation.
3 . The method of claim 2 , wherein the first modality comprises motion, and wherein the initial input comprises a gesture.
4 . The method of claim 2 , wherein the first modality comprises orientation, and wherein the initial input comprises a gaze on an object.
5 . The method of claim 1 , further comprising:
determining that the first user needs one or more suggested inputs.
6 . The method of claim 5 , wherein determining that the first user needs one or more suggested inputs is based on a wake-up input from the first user.
7 . The method of claim 6 , wherein the wake-up input comprises one or more of a voice utterance, a character string, an image, a video clip, a gesture, or a gaze.
8 . The method of claim 5 , wherein the initial input comprises a gaze on an object, and wherein determining that the first user needs one or more suggested inputs is further based on the gaze on the object.
9 . The method of claim 8 , wherein generating the one or more candidate continuation-inputs is further based on the object.
10 . The method of claim 5 , wherein determining that the first user needs one or more suggested inputs is further based on contextual information associated with the initial input.
11 . The method of claim 5 , wherein determining that the first user needs one or more suggested inputs is further based on the one or more intents.
12 . The method of claim 1 , further comprising:
identifying one or more entities associated with the one or more intents.
13 . The method of claim 12 , wherein generating the one or more candidate continuation-inputs is further based the one or more entities.
14 . The method of claim 1 , further comprising:
receiving, at the client system, a user-selected input from the first user, wherein the user-selected input comprises one of the suggested inputs; and executing one or more tasks based on the user-selected input.
15 . The method of claim 1 , further comprising:
receiving, at the client system, a first user-selected input from the first user, wherein the first user-selected input comprises one of the suggested inputs, and wherein the first user-selected input is associated with a first intent; generating, based on the first user-selected input, one or more additional candidate continuation-inputs, wherein each of the one or more additional candidate continuation-inputs is associated with the first intent; presenting, at the client system, one or more additional suggested inputs corresponding to one or more of the additional candidate continuation-inputs; receiving, at the client system, a second user-selected input from the first user, wherein the second user-selected input comprises one of the additional suggested inputs; and executing one or more tasks based on the second user-selected input.
16 . The method of claim 1 , further comprising:
determining the initial input comprises an incomplete input for triggering an execution of one or more tasks corresponding to the one or more intents.
17 . The method of claim 16 , wherein each of the candidate continuation-inputs comprises a complete instruction to trigger the execution of a respective task of the one or more tasks.
18 . The method of claim 17 , wherein each of the presented suggested inputs comprises a guidance to the first user for executing the complete instruction corresponding to the respective candidate continuation-input.
19 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:
receive, at a client system, an initial input from a first user, wherein the initial input is in a first modality; determine one or more intents and one or more slots corresponding to the initial input, wherein the one or more slots are conditioned on the one or more intents; generate, based on the one or more intents and the one or more slots, one or more candidate continuation-inputs, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, wherein the candidate modalities are different from the first modality, and wherein each of the candidate continuation-inputs references one or more entities represented by the one or more slots; and present, at the client system, one or more suggested inputs corresponding to one or more of the candidate continuation-inputs.
20 . A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:
receive, at a client system, an initial input from a first user, wherein the initial input is in a first modality; determine one or more intents and one or more slots corresponding to the initial input, wherein the one or more slots are conditioned on the one or more intents; generate, based on the one or more intents and the one or more slots, one or more candidate continuation-inputs, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, wherein the candidate modalities are different from the first modality, and wherein each of the candidate continuation-inputs references one or more entities represented by the one or more slots; and present, at the client system, one or more suggested inputs corresponding to one or more of the candidate continuation-inputs.Join the waitlist — get patent alerts
Track US2021343286A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.