US10839810B2ActiveUtilityA1

Speaker enrollment

Assignee: CIRRUS LOGIC INT SEMICONDUCTOR LTDPriority: Nov 21, 2017Filed: Nov 16, 2018Granted: Nov 17, 2020
Est. expiryNov 21, 2037(~11.3 yrs left)· nominal 20-yr term from priority
Inventors:Rahim Saeidi
G10L 17/20G10L 17/00G10L 2021/03646G10L 17/04G10L 17/02
44
PatentIndex Score
0
Cited by
12
References
13
Claims

Abstract

A method of speaker modelling for a speaker recognition system, comprises: receiving a signal comprising a speaker's speech; and, for a plurality of frames of the signal: obtaining a spectrum of the speaker's speech; generating at least one modified spectrum, by applying effects related to a respective vocal effort; and extracting features from the spectrum of the speaker's speech and the at least one modified spectrum. The method further comprises forming at least one speech model based on the extracted features.

Claims

exact text as granted — not AI-modified
The invention claimed is: 
     
       1. A method of speaker modelling for a speaker recognition system, comprising:
 receiving a signal comprising a speaker's speech; and, 
 for a plurality of frames of the signal:
 obtaining a spectrum of the speaker's speech; 
 generating at least one modified spectrum, by applying effects related to a respective vocal effort, wherein the step of generating at least one modified spectrum comprises:
 determining a frequency and a bandwidth of at least one formant component of the speaker's speech; 
 generating at least one modified formant component by modifying at least one of the frequency and the bandwidth of the or each formant component; and 
 generating the modified spectrum from the or each modified formant component; and 
 
 extracting features from the spectrum of the speaker's speech and the at least one modified spectrum; and 
 
 forming at least one speech model based on the extracted features. 
 
     
     
       2. A method according to  claim 1 , comprising:
 obtaining the spectrum of the speaker's speech for a plurality of frames of the signal containing voiced speech. 
 
     
     
       3. A method according to  claim 1 , comprising:
 obtaining the spectrum of the speaker's speech for a plurality of overlapping frames of the signal. 
 
     
     
       4. A method according to  claim 1 , wherein each frame has a duration between 10 ms and 50 ms. 
     
     
       5. A method according to  claim 1 , comprising:
 generating a plurality of modified spectra, by applying effects related to respective vocal efforts. 
 
     
     
       6. A method according to  claim 1 , wherein the step of forming at least one speech model comprises forming a background model for the speaker recognition system, based in part on said speaker's speech. 
     
     
       7. A method according to  claim 1 , comprising determining a frequency and a bandwidth of a number of formant components of the speaker's speech in the range from 3-5. 
     
     
       8. A method according to  claim 1 , wherein generating modified formant components comprises:
 modifying the frequency and the bandwidth of the or each formant component. 
 
     
     
       9. A method according to  claim 1 , wherein the features extracted from the spectrum of the user's speech comprise Mel Frequency Cepstral Coefficients. 
     
     
       10. A method according to  claim 1 , wherein the step of forming at least one speech model comprises forming a model of the speaker's speech. 
     
     
       11. A method according to  claim 10 , wherein the method is performed on enrolling the speaker in the speaker recognition system. 
     
     
       12. A non-transitory computer readable storage medium having computer-executable instructions stored thereon that, when executed by processor circuitry, cause the processor circuitry to perform a method comprising:
 receiving a signal comprising a speaker's speech; and 
 for a plurality of frames of the signal:
 obtaining a spectrum of the speaker's speech; 
 generating at least one modified spectrum, by applying effects related to a respective vocal effort, wherein the step of generating at least one modified spectrum comprises:
 determining a frequency and a bandwidth of at least one formant component of the speaker's speech; 
 generating at least one modified formant component by modifying at least one of the frequency and the bandwidth of the or each formant component; and 
 generating the modified spectrum from the or each modified formant component; 
 
 extracting features from the spectrum of the speaker's speech and the at least one modified spectrum; and 
 
 further comprising: 
 forming at least one speech model based on the extracted features. 
 
     
     
       13. A system for speaker modelling, the system comprising:
 an input, for receiving a signal comprising a speaker's speech; and, 
 a processor, configured for, for a plurality of frames of the signal:
 obtaining a spectrum of the speaker's speech; 
 generating at least one modified spectrum, by applying effects related to a respective vocal effort, wherein the step of generating at least one modified spectrum comprises:
 determining a frequency and a bandwidth of at least one formant component of the speaker's speech; 
 generating at least one modified formant component by modifying at least one of the frequency and the bandwidth of the or each formant component; and 
 generating the modified spectrum from the or each modified formant component; 
 
 extracting features from the spectrum of the speaker's speech and the at least one modified spectrum; and 
 forming at least one speech model based on the extracted features.

Join the waitlist — get patent alerts

Track US10839810B2 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.