US2023111509A1PendingUtilityA1

Detecting a trigger of a digital assistant

Assignee: APPLE INCPriority: May 16, 2017Filed: Dec 13, 2022Published: Apr 13, 2023
Est. expiryMay 16, 2037(~10.8 yrs left)· nominal 20-yr term from priority
G10L 15/22G10L 15/30G10L 15/1822H04R 2499/11H04R 2227/003G10L 2021/02166H04R 27/00G10L 2015/088G10L 2015/228G10L 15/08G10L 15/18G10L 25/51G10L 21/0216H04R 3/005H04R 3/00H04R 1/406G10L 15/04G10L 15/28
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and processes for operating an intelligent automated assistant are provided. In accordance with one example, a method includes, at an electronic device with one or more processors, memory, and a plurality of microphones, sampling, at each of the plurality of microphones of the electronic device, an audio signal to obtain a plurality of audio signals; processing the plurality of audio signals to obtain a plurality of audio streams; and determining, based on the plurality of audio streams, whether any of the plurality of audio signals corresponds to a spoken trigger. The method further includes, in accordance with a determination that the plurality of audio signals corresponds to the spoken trigger, initiating a session of the digital assistant; and in accordance with a determination that the plurality of audio signals does not correspond to the spoken trigger, foregoing initiating a session of the digital assistant.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . One or more non-transitory computer-readable storage media storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of one or more electronic devices, cause the one or more electronic devices to:
 sample, using a first microphone of a first electronic device, a first audio signal;   sample, using a second microphone of a second electronic device different from the first electronic device, a second audio signal;   determine, at the second electronic device, whether any of the first audio signal and the second audio signal corresponds to a spoken trigger, wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
 processing the first audio signal and the second audio signal to obtain a sequence of words; and 
 determining whether the sequence of words is semantically meaningful; 
   in accordance with a determination that any of the first audio signal and the second audio signal correspond to the spoken trigger:
 initiate, by the second electronic device, a session of a digital assistant; and 
   in accordance with a determination that each of the first audio signal and the second audio signal does not correspond to the spoken trigger:
 forgo, by the second electronic device, initiating the session of the digital assistant. 
   
     
     
         2 . The one or more non-transitory computer-readable storage media of  claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the one or more electronic devices to:
 obtain, by the second electronic device and based on the first audio signal and the second audio signal, directional information associated with an audio source.   
     
     
         3 . The one or more non-transitory computer-readable storage media of  claim 2 , wherein initiating the session of the digital assistant includes providing an audio output based on the directional information. 
     
     
         4 . The one or more non-transitory computer-readable storage media of  claim 3 , wherein the second electronic device includes a plurality of speakers, and wherein providing the audio output based on the directional information includes:
 providing the audio output using a first speaker of the plurality of speakers, wherein the first speaker faces the audio source; and   forgoing providing the audio output using a second speaker of the plurality of speakers.   
     
     
         5 . The one or more non-transitory computer-readable storage media of  claim 2 , wherein the second electronic device includes a plurality of microphones, and wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the one or more electronic devices to:
 select, by the second electronic device and based on the directional information, one or more microphones from the plurality of microphones; and   sample, by the second electronic device and using the selected one or more microphones, a subsequent audio signal.   
     
     
         6 . The one or more non-transitory computer-readable storage media of  claim 1 , wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
 determining whether the sequence of words is associated with a same audio source.   
     
     
         7 . The one or more non-transitory computer-readable storage media of  claim 1 , wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
 determining whether the sequence of words is indicative of an authorized user of the second electronic device.   
     
     
         8 . The one or more non-transitory computer-readable storage media of  claim 1 , wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
 determining whether the sequence of words corresponds to an actionable intent recognized by the digital assistant.   
     
     
         9 . The one or more non-transitory computer-readable storage media of  claim 1 , wherein the first electronic device is of a different device type than the second electronic device. 
     
     
         10 . The one or more non-transitory computer-readable storage media of  claim 1 , wherein the second electronic device has more processing power than the first electronic device. 
     
     
         11 . A method for operating a digital assistant, the method comprising:
 sampling, using a first microphone of a first electronic device, a first audio signal;   sampling, using a second microphone of a second electronic device different from the first electronic device, a second audio signal;   determining, at the second electronic device, whether any of the first audio signal and the second audio signal corresponds to a spoken trigger, wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
 processing the first audio signal and the second audio signal to obtain a sequence of words; and 
 determining whether the sequence of words is semantically meaningful; 
   in accordance with a determination that any of the first audio signal and the second audio signal correspond to the spoken trigger:
 initiating, by the second electronic device, a session of the digital assistant; and 
   in accordance with a determination that each of the first audio signal and the second audio signal does not correspond to the spoken trigger:
 forgoing, by the second electronic device, initiating the session of the digital assistant. 
   
     
     
         12 . A system, comprising:
 one or more processors of one or more electronic devices;   one or more memories of the one or more electronic devices; and   one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, the one or more programs including instructions for:
 sampling, using a first microphone of a first electronic device, a first audio signal; 
 sampling, using a second microphone of a second electronic device different from the first electronic device, a second audio signal; 
 determining, at the second electronic device, whether any of the first audio signal and the second audio signal corresponds to a spoken trigger, wherein determining whether any of the first audio signal and the second audio signal corresponds to the spoken trigger comprises:
 processing the first audio signal and the second audio signal to obtain a sequence of words; and 
 determining whether the sequence of words is semantically meaningful; 
 
 in accordance with a determination that any of the first audio signal and the second audio signal correspond to the spoken trigger:
 initiating, by the second electronic device, a session of the digital assistant; and 
 
 in accordance with a determination that each of the first audio signal and the second audio signal does not correspond to the spoken trigger:
 forgoing, by the second electronic device, initiating the session of the digital assistant.

Join the waitlist — get patent alerts

Track US2023111509A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.