US2008015858A1PendingUtilityA1

Methods and apparatus to perform speech reference enrollment

Assignee: BOSSEMEYER ROBERT W JRPriority: May 27, 1997Filed: Jul 9, 2007Published: Jan 17, 2008
Est. expiryMay 27, 2017(expired)· nominal 20-yr term from priority
G10L 2015/0638G10L 2015/0636H04M 3/382G10L 17/04H04M 3/493G10L 2015/0631H04M 2201/40G10L 15/07
50
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A speech reference enrollment method involves the following steps: (a) requesting a user speak a vocabulary word; (b) detecting a first utterance ( 354 ); (c) requesting the user speak the vocabulary word; (d) detecting a second utterance ( 358 ); (e) determining a first similarity between the first utterance and the second utterance ( 362 ); (f) when the first similarity is less than a predetermined similarity, requesting the user speak the vocabulary word; (g) detecting a third utterance ( 366 ); (h) determining a second similarity between the first utterance and the third utterance ( 370 ); and (i) when the second similarity is greater than or equal to the predetermined similarity, creating a reference ( 364 ).

Claims

exact text as granted — not AI-modified
1 - 22 . (canceled)  
   
   
       23 . A method, comprising: 
 receiving a first utterance of a word;    receiving a second utterance of the word;    when a number of voiced speech frames associated with the second utterance is greater than a threshold, determining a first similarity between the first utterance and the second utterance; and    when the first similarity is greater than or equal to a similarity threshold, storing a reference for the word.    
   
   
       24 . A method as defined in  claim 23 , wherein determining the first similarity between the first utterance and the second utterance comprises determining a similarity between a first plurality of features associated with the first utterance and a second plurality of features associated with the second utterance.  
   
   
       25 . A method as defined in  claim 23 , further comprising: 
 when the first similarity is less than the similarity threshold, requesting a user to speak a third utterance of the word;    determining a second similarity between the first utterance and the third utterance; and    when the second similarity is greater than or equal to the similarity threshold, storing the reference for the word.    
   
   
       26 . A method as defined in  claim 23 , further comprising determining the number of voiced speech frames by estimating the number of voiced speech frames.  
   
   
       27 . A method as defined in  claim 23 , further comprising requesting the user repeat the word when the number of voiced speech frames is less than the threshold.  
   
   
       28 . A method as defined in  claim 23 , further comprising: 
 determining a signal to noise ratio of the first utterance; and    when the signal to noise ratio is less than a threshold signal to noise ratio, increasing a gain of a voice amplifier.    
   
   
       29 . A method as defined in  claim 23 , further comprising determining an amplitude histogram of the first utterance.  
   
   
       30 . A method as defined in  claim 23 , further comprising retrieving the reference and verifying a speaker based on the reference.  
   
   
       31 . A method, comprising: 
 receiving a first utterance of a word;    receiving a second utterance of the word;    when the duration of the second utterance is greater than or equal to a first duration or less than or equal to a second duration, storing a reference for the word.    
   
   
       32 . A method as defined in  claim 31 , further comprising: 
 when the duration of the second utterance is greater than or equal to the first duration or less than or equal to the second duration, determining a first similarity between the first utterance and the second utterance; and    when the first similarity is greater than or equal to a similarity threshold, storing the reference.    
   
   
       33 . A method as defined in  claim 32 , further comprising: 
 when the first similarity is less than the similarity threshold, requesting the user speak the word;    determining a second similarity between the first utterance and a third utterance; and    when the second similarity is greater than or equal to the similarity threshold, storing the reference.    
   
   
       34 . A method as defined in  claim 33 , further comprising: 
 determining a third similarity between the second utterance and the third utterance; and    when the third similarity is greater than or equal to the similarity threshold, storing the reference.    
   
   
       35 . A method as defined in  claim 31 , further comprising: 
 determining if the first utterance exceeds an amplitude threshold within a time period; and    when the first utterance does not exceed the amplitude threshold within the time period, requesting the user re-speak the word.    
   
   
       36 . A method as defined in  claim 31 , further comprising: 
 associating a start time with a point at which a first amplitude of the first utterance is greater than an amplitude threshold;    associating an end time with a point at which a second amplitude of the first utterance is less than the amplitude threshold; and    determining the duration as a difference between the end time and the start time.    
   
   
       37 . A method as defined in  claim 35 , further comprising: 
 determining a signal to noise ratio of the first utterance; and    when the signal to noise ratio is less than a threshold signal to noise ratio, increasing a gain of a voice amplifier.    
   
   
       38 . A method as defined in  claim 31 , further comprising retrieving the reference and verifying a speaker based on the reference.  
   
   
       39 . A machine accessible storage medium having instructions stored thereon that, when executed, cause a machine to: 
 receive a first utterance of a word;    receive a second utterance of the word;    when the duration of the second utterance is greater than or equal to a first duration or less than or equal to a second duration, store a reference for the word.    
   
   
       40 . A machine accessible storage medium as defined in  claim 39  having instructions stored thereon that, when executed, cause the machine to: 
 when the duration of the second utterance is greater than or equal to the first duration or less than or equal to the second duration:    determine a first similarity between the first utterance and the second utterance; and    when the first similarity is greater than or equal to a similarity threshold, store the reference.    
   
   
       41 . A machine accessible storage medium as defined in  claim 40  having instructions stored thereon that, when executed, cause the machine to: 
 when the first similarity is less than the similarity threshold:    request the user speak the word;    determine a second similarity between the first utterance and a third utterance; and    when the second similarity is greater than or equal to the similarity threshold, store the reference.    
   
   
       42 . A machine accessible storage medium as defined in  claim 41  having instructions stored thereon that, when executed, cause the machine to: 
 determine a third similarity between the second utterance and the third utterance; and    when the third similarity is greater than or equal to the predetermined similarity, store the reference.    
   
   
       43 . A machine accessible storage medium as defined in  claim 39  having instructions stored thereon that, when executed, cause the machine to: 
 determine if the first utterance exceeds an amplitude threshold within a time period; and    when the first utterance does not exceed the amplitude threshold within the time period, request the user re-speak the word.    
   
   
       44 . A machine accessible storage medium as defined in  claim 39  having instructions stored thereon that, when executed, cause the machine to: 
 associate a start time with a point at which a first amplitude of the first utterance is greater than an amplitude threshold;    associate an end time with a point at which a second amplitude of the first utterance is less than the amplitude threshold; and    determine the duration as a difference between the end time and the start time.    
   
   
       45 . A machine accessible storage medium as defined in  claim 39  having instructions stored thereon that, when executed, cause the machine to: 
 determine a signal to noise ratio of the first utterance; and    when the signal to noise ratio is less than a threshold signal to noise ratio, increase a gain of a voice amplifier.    
   
   
       46 . A machine accessible storage medium as defined in  claim 39  having instructions stored thereon that, when executed, cause the machine to determine an amplitude histogram of the first utterance.  
   
   
       47 . A machine accessible storage medium as defined in  claim 39  having instructions stored thereon that, when executed, cause the machine to retrieve the reference and verify a speaker based on the reference.

Join the waitlist — get patent alerts

Track US2008015858A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.