US10839810B2ActiveUtilityA1
Speaker enrollment
Assignee: CIRRUS LOGIC INT SEMICONDUCTOR LTDPriority: Nov 21, 2017Filed: Nov 16, 2018Granted: Nov 17, 2020
Est. expiryNov 21, 2037(~11.3 yrs left)· nominal 20-yr term from priority
Inventors:Rahim Saeidi
G10L 17/20G10L 17/00G10L 2021/03646G10L 17/04G10L 17/02
44
PatentIndex Score
0
Cited by
12
References
13
Claims
Abstract
A method of speaker modelling for a speaker recognition system, comprises: receiving a signal comprising a speaker's speech; and, for a plurality of frames of the signal: obtaining a spectrum of the speaker's speech; generating at least one modified spectrum, by applying effects related to a respective vocal effort; and extracting features from the spectrum of the speaker's speech and the at least one modified spectrum. The method further comprises forming at least one speech model based on the extracted features.
Claims
exact text as granted — not AI-modifiedThe invention claimed is:
1. A method of speaker modelling for a speaker recognition system, comprising:
receiving a signal comprising a speaker's speech; and,
for a plurality of frames of the signal:
obtaining a spectrum of the speaker's speech;
generating at least one modified spectrum, by applying effects related to a respective vocal effort, wherein the step of generating at least one modified spectrum comprises:
determining a frequency and a bandwidth of at least one formant component of the speaker's speech;
generating at least one modified formant component by modifying at least one of the frequency and the bandwidth of the or each formant component; and
generating the modified spectrum from the or each modified formant component; and
extracting features from the spectrum of the speaker's speech and the at least one modified spectrum; and
forming at least one speech model based on the extracted features.
2. A method according to claim 1 , comprising:
obtaining the spectrum of the speaker's speech for a plurality of frames of the signal containing voiced speech.
3. A method according to claim 1 , comprising:
obtaining the spectrum of the speaker's speech for a plurality of overlapping frames of the signal.
4. A method according to claim 1 , wherein each frame has a duration between 10 ms and 50 ms.
5. A method according to claim 1 , comprising:
generating a plurality of modified spectra, by applying effects related to respective vocal efforts.
6. A method according to claim 1 , wherein the step of forming at least one speech model comprises forming a background model for the speaker recognition system, based in part on said speaker's speech.
7. A method according to claim 1 , comprising determining a frequency and a bandwidth of a number of formant components of the speaker's speech in the range from 3-5.
8. A method according to claim 1 , wherein generating modified formant components comprises:
modifying the frequency and the bandwidth of the or each formant component.
9. A method according to claim 1 , wherein the features extracted from the spectrum of the user's speech comprise Mel Frequency Cepstral Coefficients.
10. A method according to claim 1 , wherein the step of forming at least one speech model comprises forming a model of the speaker's speech.
11. A method according to claim 10 , wherein the method is performed on enrolling the speaker in the speaker recognition system.
12. A non-transitory computer readable storage medium having computer-executable instructions stored thereon that, when executed by processor circuitry, cause the processor circuitry to perform a method comprising:
receiving a signal comprising a speaker's speech; and
for a plurality of frames of the signal:
obtaining a spectrum of the speaker's speech;
generating at least one modified spectrum, by applying effects related to a respective vocal effort, wherein the step of generating at least one modified spectrum comprises:
determining a frequency and a bandwidth of at least one formant component of the speaker's speech;
generating at least one modified formant component by modifying at least one of the frequency and the bandwidth of the or each formant component; and
generating the modified spectrum from the or each modified formant component;
extracting features from the spectrum of the speaker's speech and the at least one modified spectrum; and
further comprising:
forming at least one speech model based on the extracted features.
13. A system for speaker modelling, the system comprising:
an input, for receiving a signal comprising a speaker's speech; and,
a processor, configured for, for a plurality of frames of the signal:
obtaining a spectrum of the speaker's speech;
generating at least one modified spectrum, by applying effects related to a respective vocal effort, wherein the step of generating at least one modified spectrum comprises:
determining a frequency and a bandwidth of at least one formant component of the speaker's speech;
generating at least one modified formant component by modifying at least one of the frequency and the bandwidth of the or each formant component; and
generating the modified spectrum from the or each modified formant component;
extracting features from the spectrum of the speaker's speech and the at least one modified spectrum; and
forming at least one speech model based on the extracted features.Join the waitlist — get patent alerts
Track US10839810B2 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.