Methods and apparatus to perform speech reference enrollment
Abstract
A speech reference enrollment method involves the following steps: (a) requesting a user speak a vocabulary word; (b) detecting a first utterance ( 354 ); (c) requesting the user speak the vocabulary word; (d) detecting a second utterance ( 358 ); (e) determining a first similarity between the first utterance and the second utterance ( 362 ); (f) when the first similarity is less than a predetermined similarity, requesting the user speak the vocabulary word; (g) detecting a third utterance ( 366 ); (h) determining a second similarity between the first utterance and the third utterance ( 370 ); and (i) when the second similarity is greater than or equal to the predetermined similarity, creating a reference ( 364 ).
Claims
exact text as granted — not AI-modified1 - 22 . (canceled)
23 . A method, comprising:
receiving a first utterance of a word; receiving a second utterance of the word; when a number of voiced speech frames associated with the second utterance is greater than a threshold, determining a first similarity between the first utterance and the second utterance; and when the first similarity is greater than or equal to a similarity threshold, storing a reference for the word.
24 . A method as defined in claim 23 , wherein determining the first similarity between the first utterance and the second utterance comprises determining a similarity between a first plurality of features associated with the first utterance and a second plurality of features associated with the second utterance.
25 . A method as defined in claim 23 , further comprising:
when the first similarity is less than the similarity threshold, requesting a user to speak a third utterance of the word; determining a second similarity between the first utterance and the third utterance; and when the second similarity is greater than or equal to the similarity threshold, storing the reference for the word.
26 . A method as defined in claim 23 , further comprising determining the number of voiced speech frames by estimating the number of voiced speech frames.
27 . A method as defined in claim 23 , further comprising requesting the user repeat the word when the number of voiced speech frames is less than the threshold.
28 . A method as defined in claim 23 , further comprising:
determining a signal to noise ratio of the first utterance; and when the signal to noise ratio is less than a threshold signal to noise ratio, increasing a gain of a voice amplifier.
29 . A method as defined in claim 23 , further comprising determining an amplitude histogram of the first utterance.
30 . A method as defined in claim 23 , further comprising retrieving the reference and verifying a speaker based on the reference.
31 . A method, comprising:
receiving a first utterance of a word; receiving a second utterance of the word; when the duration of the second utterance is greater than or equal to a first duration or less than or equal to a second duration, storing a reference for the word.
32 . A method as defined in claim 31 , further comprising:
when the duration of the second utterance is greater than or equal to the first duration or less than or equal to the second duration, determining a first similarity between the first utterance and the second utterance; and when the first similarity is greater than or equal to a similarity threshold, storing the reference.
33 . A method as defined in claim 32 , further comprising:
when the first similarity is less than the similarity threshold, requesting the user speak the word; determining a second similarity between the first utterance and a third utterance; and when the second similarity is greater than or equal to the similarity threshold, storing the reference.
34 . A method as defined in claim 33 , further comprising:
determining a third similarity between the second utterance and the third utterance; and when the third similarity is greater than or equal to the similarity threshold, storing the reference.
35 . A method as defined in claim 31 , further comprising:
determining if the first utterance exceeds an amplitude threshold within a time period; and when the first utterance does not exceed the amplitude threshold within the time period, requesting the user re-speak the word.
36 . A method as defined in claim 31 , further comprising:
associating a start time with a point at which a first amplitude of the first utterance is greater than an amplitude threshold; associating an end time with a point at which a second amplitude of the first utterance is less than the amplitude threshold; and determining the duration as a difference between the end time and the start time.
37 . A method as defined in claim 35 , further comprising:
determining a signal to noise ratio of the first utterance; and when the signal to noise ratio is less than a threshold signal to noise ratio, increasing a gain of a voice amplifier.
38 . A method as defined in claim 31 , further comprising retrieving the reference and verifying a speaker based on the reference.
39 . A machine accessible storage medium having instructions stored thereon that, when executed, cause a machine to:
receive a first utterance of a word; receive a second utterance of the word; when the duration of the second utterance is greater than or equal to a first duration or less than or equal to a second duration, store a reference for the word.
40 . A machine accessible storage medium as defined in claim 39 having instructions stored thereon that, when executed, cause the machine to:
when the duration of the second utterance is greater than or equal to the first duration or less than or equal to the second duration: determine a first similarity between the first utterance and the second utterance; and when the first similarity is greater than or equal to a similarity threshold, store the reference.
41 . A machine accessible storage medium as defined in claim 40 having instructions stored thereon that, when executed, cause the machine to:
when the first similarity is less than the similarity threshold: request the user speak the word; determine a second similarity between the first utterance and a third utterance; and when the second similarity is greater than or equal to the similarity threshold, store the reference.
42 . A machine accessible storage medium as defined in claim 41 having instructions stored thereon that, when executed, cause the machine to:
determine a third similarity between the second utterance and the third utterance; and when the third similarity is greater than or equal to the predetermined similarity, store the reference.
43 . A machine accessible storage medium as defined in claim 39 having instructions stored thereon that, when executed, cause the machine to:
determine if the first utterance exceeds an amplitude threshold within a time period; and when the first utterance does not exceed the amplitude threshold within the time period, request the user re-speak the word.
44 . A machine accessible storage medium as defined in claim 39 having instructions stored thereon that, when executed, cause the machine to:
associate a start time with a point at which a first amplitude of the first utterance is greater than an amplitude threshold; associate an end time with a point at which a second amplitude of the first utterance is less than the amplitude threshold; and determine the duration as a difference between the end time and the start time.
45 . A machine accessible storage medium as defined in claim 39 having instructions stored thereon that, when executed, cause the machine to:
determine a signal to noise ratio of the first utterance; and when the signal to noise ratio is less than a threshold signal to noise ratio, increase a gain of a voice amplifier.
46 . A machine accessible storage medium as defined in claim 39 having instructions stored thereon that, when executed, cause the machine to determine an amplitude histogram of the first utterance.
47 . A machine accessible storage medium as defined in claim 39 having instructions stored thereon that, when executed, cause the machine to retrieve the reference and verify a speaker based on the reference.Join the waitlist — get patent alerts
Track US2008015858A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.