US2008059168A1PendingUtilityA1

Speech recognition using discriminant features

Assignee: IBMPriority: Mar 29, 2001Filed: Oct 31, 2007Published: Mar 6, 2008
Est. expiryMar 29, 2021(expired)· nominal 20-yr term from priority
Inventors:Ellen M. Eide
G10L 15/02
50
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods and arrangements for representing the speech waveform in terms of a set of abstract, linguistic distinctions in order to derive a set of discriminative features for use in a speech recognizer. By combining the distinctive feature representation with an original waveform representation, it is possible to achieve a reduction in word error rate of 33% on an automatic speech recognition task.

Claims

exact text as granted — not AI-modified
1 . A method of facilitating speech recognition, said method comprising the steps of: 
 obtaining speech input data;    building a model for each feature of an original set of linguistic features, wherein the model reflects whether or not each feature is present;    ranking the linguistic features; and    rebuilding the model for each of a preselected number N of the ranked linguistic features.    
     
     
         2 . The method according to  claim 1 , wherein said step of building a model for each of a preselected number N of the ranked features comprises building a model for the top N ranked features.  
     
     
         3 . The method according to  claim 1 , further comprising the step of compiling a confusion matrix for each feature of the original set of features subsequent to said step of building a model for each feature of an original set of features.  
     
     
         4 . The method according to  claim 3 , wherein said step of compiling a confusion matrix comprises computing a score for each feature based on the likelihood of its presence in a frame of the speech input data.  
     
     
         5 . The method according to  claim 4 , wherein said step of computing a score for each feature comprises computing a score as a log-likelihood ratio.  
     
     
         6 . The method according to  claim 4 , wherein said step of compiling a confusion matrix further comprises comparing each score of each feature with a threshold.  
     
     
         7 . The method according to  claim 4 , wherein said step of compiling a confusion matrix further comprises calculating mutual information between truth and labels for each feature.  
     
     
         8 . The method according to  claim 1 , wherein said step of building a model for each feature of an original set of features comprises: 
 partitioning the speech input data in parallel, once for each feature; and    producing an observation vector.    
     
     
         9 . The method according to  claim 8 , wherein said step of building a model for each feature of an original set of features comprises: 
 partitioning data in parallel from the observation vector, once for each feature; and    producing final observations.    
     
     
         10 . The method according to  claim 1 , wherein said step of building a model for each of a preselected number N of the ranked features comprises: 
 partitioning the speech input data in parallel, once for each feature; and    producing an observation vector.    
     
     
         11 . The method according to  claim 10 , wherein said step of building a model for each of a preselected number N of the ranked features comprises: 
 partitioning data in parallel from the observation vector, once for each feature; and    producing final observations.    
     
     
         12 . An apparatus for facilitating speech recognition, said method comprising the steps of: 
 an input medium which obtains speech input data;    a first model builder which builds a model for each feature of an original set of linguistic features, wherein the model reflects whether or not each feature is present;    a ranking arrangement which ranks the linguistic features; and    a second model builder which rebuilds the model for each of a preselected number N of the ranked linguistic features.    
     
     
         13 . The apparatus according to  claim 12 , wherein said second model builder is adapted to build a model for the top N ranked features.  
     
     
         14 . The apparatus according to  claim 12 , further comprising a matrix compiler which compiles a confusion matrix for each feature of the original set of features subsequent to said step of building a model for each feature of an original set of features.  
     
     
         15 . The apparatus according to  claim 14 , wherein said matrix compiler is adapted to compute a score for each feature based on the likelihood of its presence in a frame of the speech input data.  
     
     
         16 . The apparatus according to  claim 15 , wherein said matrix compiler is adapted to compute a score as a log-likelihood ratio.  
     
     
         17 . The apparatus according to  claim 15 , wherein said matrix compiler is adapted to compare each score of each feature with a threshold.  
     
     
         18 . The apparatus according to  claim 15 , wherein said matrix compiler is adapted to calculate mutual information between truth and labels for each feature..  
     
     
         19 . The apparatus according to  claim 12 , wherein said first model builder is adapted to: 
 partition the speech input data in parallel, once for each feature; and    produce an observation vector.    
     
     
         20 . The apparatus according to  claim 19 , wherein said first model builder is adapted to: 
 partition data in parallel from the observation vector, once for each feature; and    produce final observations.    
     
     
         21 . The apparatus according to  claim 12 , wherein said second model builder is adapted to: 
 partition the speech input data in parallel, once for each feature; and    produce an observation vector.    
     
     
         22 . The apparatus according to  claim 21 , wherein said second model builder is adapted to: 
 partition data in parallel from the observation vector, once for each feature; and    produce final observations.    
     
     
         23 . A program storage device readable by machine, tangibly embodying a program of instructions executable by the machine to perform method steps for speech recognition, said method comprising the steps of: 
 obtaining speech input data;    building a model for each feature of an original set of linguistic features, wherein the model reflects whether or not each feature is present;    ranking the linguistic features; and    rebuilding the model for each of a preselected number N of the ranked linguistic features.

Join the waitlist — get patent alerts

Track US2008059168A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.