Detecting a trigger of a digital assistant
Abstract
Systems and processes for operating an intelligent automated assistant are provided. In accordance with one example, a method includes, at an electronic device with one or more processors, memory, and a plurality of microphones, sampling, at each of the plurality of microphones of the electronic device, an audio signal to obtain a plurality of audio signals; processing the plurality of audio signals to obtain a plurality of audio streams; and determining, based on the plurality of audio streams, whether any of the plurality of audio signals corresponds to a spoken trigger. The method further includes, in accordance with a determination that the plurality of audio signals corresponds to the spoken trigger, initiating a session of the digital assistant; and in accordance with a determination that the plurality of audio signals does not correspond to the spoken trigger, foregoing initiating a session of the digital assistant.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . One or more non-transitory computer-readable storage media storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of one or more electronic devices, cause the one or more electronic devices to:
sample, using a first microphone of a first electronic device, a first audio signal; sample, using a second microphone of a second electronic device different from the first electronic device, a second audio signal; determine, at the second electronic device, whether any of the first audio signal and the second audio signal corresponds to a spoken trigger, wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
processing the first audio signal and the second audio signal to obtain a sequence of words; and
determining whether the sequence of words is semantically meaningful;
in accordance with a determination that any of the first audio signal and the second audio signal correspond to the spoken trigger:
initiate, by the second electronic device, a session of a digital assistant; and
in accordance with a determination that each of the first audio signal and the second audio signal does not correspond to the spoken trigger:
forgo, by the second electronic device, initiating the session of the digital assistant.
2 . The one or more non-transitory computer-readable storage media of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the one or more electronic devices to:
obtain, by the second electronic device and based on the first audio signal and the second audio signal, directional information associated with an audio source.
3 . The one or more non-transitory computer-readable storage media of claim 2 , wherein initiating the session of the digital assistant includes providing an audio output based on the directional information.
4 . The one or more non-transitory computer-readable storage media of claim 3 , wherein the second electronic device includes a plurality of speakers, and wherein providing the audio output based on the directional information includes:
providing the audio output using a first speaker of the plurality of speakers, wherein the first speaker faces the audio source; and forgoing providing the audio output using a second speaker of the plurality of speakers.
5 . The one or more non-transitory computer-readable storage media of claim 2 , wherein the second electronic device includes a plurality of microphones, and wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the one or more electronic devices to:
select, by the second electronic device and based on the directional information, one or more microphones from the plurality of microphones; and sample, by the second electronic device and using the selected one or more microphones, a subsequent audio signal.
6 . The one or more non-transitory computer-readable storage media of claim 1 , wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
determining whether the sequence of words is associated with a same audio source.
7 . The one or more non-transitory computer-readable storage media of claim 1 , wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
determining whether the sequence of words is indicative of an authorized user of the second electronic device.
8 . The one or more non-transitory computer-readable storage media of claim 1 , wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
determining whether the sequence of words corresponds to an actionable intent recognized by the digital assistant.
9 . The one or more non-transitory computer-readable storage media of claim 1 , wherein the first electronic device is of a different device type than the second electronic device.
10 . The one or more non-transitory computer-readable storage media of claim 1 , wherein the second electronic device has more processing power than the first electronic device.
11 . A method for operating a digital assistant, the method comprising:
sampling, using a first microphone of a first electronic device, a first audio signal; sampling, using a second microphone of a second electronic device different from the first electronic device, a second audio signal; determining, at the second electronic device, whether any of the first audio signal and the second audio signal corresponds to a spoken trigger, wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
processing the first audio signal and the second audio signal to obtain a sequence of words; and
determining whether the sequence of words is semantically meaningful;
in accordance with a determination that any of the first audio signal and the second audio signal correspond to the spoken trigger:
initiating, by the second electronic device, a session of the digital assistant; and
in accordance with a determination that each of the first audio signal and the second audio signal does not correspond to the spoken trigger:
forgoing, by the second electronic device, initiating the session of the digital assistant.
12 . A system, comprising:
one or more processors of one or more electronic devices; one or more memories of the one or more electronic devices; and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, the one or more programs including instructions for:
sampling, using a first microphone of a first electronic device, a first audio signal;
sampling, using a second microphone of a second electronic device different from the first electronic device, a second audio signal;
determining, at the second electronic device, whether any of the first audio signal and the second audio signal corresponds to a spoken trigger, wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
processing the first audio signal and the second audio signal to obtain a sequence of words; and
determining whether the sequence of words is semantically meaningful;
in accordance with a determination that any of the first audio signal and the second audio signal correspond to the spoken trigger:
initiating, by the second electronic device, a session of the digital assistant; and
in accordance with a determination that each of the first audio signal and the second audio signal does not correspond to the spoken trigger:
forgoing, by the second electronic device, initiating the session of the digital assistant.Join the waitlist — get patent alerts
Track US2023111509A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.