US2015046157A1PendingUtilityA1
User Dedicated Automatic Speech Recognition
Est. expiryMar 16, 2032(~5.6 yrs left)· nominal 20-yr term from priority
G06F 3/167G10L 25/51G10L 15/22G10L 2015/228G10L 15/28G10L 15/183G10L 2021/02166G10L 17/00G10L 15/25
45
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A multi-mode voice controlled user interface is described. The user interface is adapted to conduct a speech dialog with one or more possible speakers and includes a broad listening mode which accepts speech inputs from the possible speakers without spatial filtering, and a selective listening mode which limits speech inputs to a specific speaker using spatial filtering. The user interface switches listening modes in response to one or more switching cues.
Claims
exact text as granted — not AI-modified1 . A device for automatic speech recognition (ASR) comprising:
a multi-mode voice controlled user interface employing at least one hardware implemented computer processor, wherein the user interface is adapted to conduct a speech dialog with one or more possible speakers and includes:
a broad listening mode which accepts speech inputs from the possible speakers without spatial filtering; and
a selective listening mode which limits speech inputs to a specific speaker using spatial filtering;
wherein the user interface switches listening modes in response to one or more switching cues.
2 . A device according to claim 1 , wherein the broad listening mode uses an associated broad mode recognition vocabulary and the selective listening mode uses a different associated selective mode recognition vocabulary.
3 . A device according to claim 1 , wherein the switching cues include one or more mode switching words from the speech inputs.
4 . A device according to claim 1 , wherein the switching cues include one or more dialog states in the speech dialog.
5 . A device according to claim 1 , wherein the switching cues include one or more visual cues from the possible speakers.
6 . A device according to claim 1 , wherein the selective listening mode uses acoustic speaker localization for the spatial filtering.
7 . A device according to claim 1 , wherein the selective listening mode uses image processing for the spatial filtering.
8 . A device according to claim 1 , wherein the user interface operates in selective listening mode simultaneously in parallel for each of a plurality of selected speakers.
9 . A device according to claim 1 , wherein the interface is adapted to operate in both listening modes in parallel, whereby the interface accepts speech inputs from any user in the room in the broad listening mode, and at the same time accepts speech inputs from only one selected speaker in the selective listening mode.
10 . A computer program product encoded in a non-transitory computer-readable medium for operating an automatic speech recognition (ASR) system, the product comprising:
program code for conducting a speech dialog with one or more possible speakers via a multi-mode voice controlled user interface adapted to:
accept speech inputs from the possible speakers in a broad listening mode without spatial filtering; and
limit speech inputs to a specific speaker in a selective listening mode using spatial filtering;
wherein the user interface switches listening modes in response to one or more switching cues.
11 . (canceled)
12 . (canceled)
13 . (canceled)
14 . (canceled)
15 . (canceled)
16 . (canceled)
17 . (canceled)
18 . A method for automatic speech recognition (ASR) comprising:
employing a multi-mode voice controlled user interface having a computer processor to conduct a speech dialog with one or more possible speakers by: employing a broad listening mode which accepts speech inputs from the possible speakers without spatial filtering; and employing a selective listening mode which limits speech inputs to a specific speaker using spatial filtering; wherein the user interface switches listening modes in response to one or more switching cues.
19 . The method according to claim 18 , wherein the broad listening mode uses an associated broad mode recognition vocabulary and the selective listening mode uses a different associated selective mode recognition vocabulary.
20 . The method according to claim 18 , wherein the switching cues include one or more mode switching words from the speech inputs.
21 . The method according to claim 18 , wherein the switching cues include one or more dialog states in the speech dialog.
22 . The method according to claim 18 , wherein the switching cues include one or more visual cues from the possible speakers.
23 . The method according to claim 18 , wherein the selective listening mode includes using acoustic speaker localization for the spatial filtering.
24 . The method according to claim 18 , wherein the selective listening mode includes using image processing for the spatial filtering.
25 . The method according to claim 18 , wherein the user interface operates in selective listening mode simultaneously in parallel for each of a plurality of selected speakers.
26 . The method according to claim 18 , wherein the user interface operates in both listening modes in parallel, such that the interface accepts speech inputs from any user in the room in the broad listening mode, and at the same time accepts speech inputs from only one selected speaker in the selective listening mode.Join the waitlist — get patent alerts
Track US2015046157A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.