Speaker verification
Abstract
Methods, systems, apparatus, including computer programs encoded on computer storage medium, to facilitate language independent-speaker verification. In one aspect, a method includes actions of receiving, by a user device, audio data representing an utterance of a user. Other actions may include providing, to a neural network stored on the user device, input data derived from the audio data and a language identifier. The neural network may be trained using speech data representing speech in different languages or dialects. The method may include additional actions of generating, based on output of the neural network produced in response to receiving the set of input data, a speaker representation and determining, based on the speaker representation and a second representation, that the utterance is an utterance of the user. The method may provide the user with access to the user device based on determining that the utterance is an utterance of the user.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method comprising:
receiving, by a mobile device that implements a language-independent speaker verification model comprising a neural network that is stored on the mobile device and configured to determine whether received audio data likely includes an utterance of one of multiple language-specific hotwords, (i) particular audio data corresponding to a particular utterance of a user, and (ii) data indicating a particular language spoken by the user; and in response to receiving (i) the particular audio data corresponding to a particular utterance of a user, and (ii) the data indicating a particular language spoken by the user, providing, for output, an indication that the language-independent speaker verification model has determined that the particular audio data likely includes the utterance of a hotword designated for the particular language spoken by the user.
2 . The computer-implemented method of claim 1 , wherein providing, for output, the indication comprises providing access to a resource of the mobile device.
3 . The computer-implemented method of claim 1 , wherein providing, for output, the indication comprises unlocking the mobile device.
4 . The computer-implemented method of claim 1 , wherein providing, for output, the indication comprises waking up the mobile device from a low-power state.
5 . The computer-implemented method of claim 1 , wherein providing, for output, the indication comprises providing an indication that the language-independent speaker verification model has determined that the particular audio data includes the utterance of a particular user associated with the mobile device.
6 . The computer-implemented method of claim 1 , wherein the neural network of the language-independent speaker verification model is trained without using utterances of the user.
7 . A method comprising:
receiving, by a user device, audio data representing an utterance of a user; providing, to a language independent speaker verification model comprising a neural network stored on the user device, a set of input data derived from the audio data and a language identifier or location identifier associated with the user device, the neural network having parameters trained using speech data representing speech in different languages or dialects; generating, based on output of the language independent speaker verification model produced in response to receiving the set of input data, a speaker representation indicative of characteristics of the voice of the user; determining, based on the speaker representation and a second representation, that the utterance is an utterance of the user; and providing the user access to the user device based on determining that the utterance is an utterance of the user.
8 . The method of claim 7 , wherein the set of input data derived from the audio data and the determined language identifier includes a first vector that is derived from the audio data and a second vector that is derived from a language identifier associated with the user device.
9 . The method of claim 8 , further comprising:
generating an input vector by concatenating the first vector and the second vector into a single concatenated vector; providing, to the language independent speaker verification model, the generated input vector; and generating, based on output of the language independent speaker verification model produced in response to receiving the input vector, a speaker representation indicative of characteristics of the voice of the user.
10 . The method of claim 8 , the method further comprising:
generating an input vector by concatenating the outputs of at least two other neural networks that respectively generate outputs based on (i) the first vector, (ii) the second vector, or (iii) both the first vector and the second vector; providing, to the language independent speaker verification model, the generated input vector; and generating, based on output of the language independent speaker verification model produced in response to receiving the input vector, a speaker representation indicative of characteristics of the voice of the user.
11 . The method of claim 8 , further comprising:
generating an input vector based on the weighted sum of the first vector and the second vector; providing, to the language independent speaker verification model, the generated input vector; and generating, based on output of the language independent speaker verification model produced in response to receiving the input vector, a speaker representation indicative of characteristics of the voice of the user.
12 . The method of claim 7 , wherein the output of the language independent speaker verification model produced in response to receiving the set of input data includes a set of activations generated by a hidden layer of the neural network.
13 . The method of claim 7 , wherein determining, based on the speaker representation and a second representation, that the utterance is an utterance of the user comprises:
determining a distance between the first representation and the second representation.
14 . The method of claim 7 , wherein providing the user access to the user device based on determining that the utterance is an utterance of the user includes unlocking the user device.
15 . A system comprising:
one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
receiving, by a user device, audio data representing an utterance of a user;
providing, to a language independent speaker verification model comprising a neural network stored on the user device, a set of input data derived from the audio data and a language identifier or location identifier associated with the user device, the neural network having parameters trained using speech data representing speech in different languages or different dialects;
generating, based on output of the language independent speaker verification model produced in response to receiving the set of input data, a speaker representation indicative of characteristics of the voice of the user;
determining, based on the speaker representation and a second representation, that the utterance is an utterance of the user; and
providing the user access to the user device based on determining that the utterance is an utterance of the user.
16 . The system of claim 15 , wherein the set of input data derived from the audio data and the determined language identifier includes a first vector that is derived from the audio data and a second vector that is derived from a language identifier associated with the user device.
17 . The system of claim 16 , further comprising:
generating an input vector by concatenating the first vector and the second vector into a single concatenated vector; providing, to the language independent speaker verification model, the generated input vector; and generating, based on output of the language independent speaker verification model produced in response to receiving the input vector, a speaker representation indicative of characteristics of the voice of the user.
18 . The system of claim 16 , the method further comprising:
generating an input vector by concatenating the outputs of at least two other neural networks that respectively generate outputs based on (i) the first vector, (ii) the second vector, or (iii) both the first vector and the second vector; providing, to the language independent speaker verification model, the generated input vector; and generating, based on output of the language independent speaker verification model produced in response to receiving the input vector, a speaker representation indicative of characteristics of the voice of the user.
19 . The system of claim 16 , further comprising:
generating an input vector based on the weighted sum of the first vector and the second vector; providing, to the language independent speaker verification model, the generated input vector; and generating, based on output of the language independent speaker verification model produced in response to receiving the input vector, a speaker representation indicative of characteristics of the voice of the user.
20 . The system of claim 15 , wherein the output of the language independent speaker verification model produced in response to receiving the set of input data includes a set of activations generated by a hidden layer of the neural network.Join the waitlist — get patent alerts
Track US2018018973A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.