US2020012969A1PendingUtilityA1

Model training method, apparatus, and device, and data similarity determining method, apparatus, and device

Assignee: ALIBABA GROUP HOLDING LTDPriority: Jul 19, 2017Filed: Sep 20, 2019Published: Jan 9, 2020
Est. expiryJul 19, 2037(~11 yrs left)· nominal 20-yr term from priority
G06N 20/20G06V 10/774G06V 10/761G06N 20/00G06N 7/01G06F 18/214G06F 18/22G10L 17/04G10L 15/02G10L 15/063G06K 9/6256G06K 9/00288G06K 9/00268G06V 40/16G06V 40/172G06V 40/168G06N 20/10G06V 10/765G10L 15/08
57
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A model training method includes: acquiring a plurality of user data pairs, wherein data fields of two sets of user data in each user data pair have an identical part; acquiring a user similarity corresponding to each user data pair, wherein the user similarity is a similarity between users corresponding to the two sets of user data in each user data pair; determining, according to the user similarity corresponding to each user data pair and the plurality of user data pairs, sample data for training a preset classification model; and training the classification model based on the sample data to obtain a similarity classification model.

Claims

exact text as granted — not AI-modified
1 . A model training method, comprising:
 acquiring a plurality of user data pairs, wherein data fields of two sets of user data in each user data pair have an identical part;   acquiring a user similarity corresponding to each user data pair, wherein the user similarity is a similarity between users corresponding to the two sets of user data in each user data pair;   determining, according to the user similarity corresponding to each user data pair and the plurality of user data pairs, sample data for training a preset classification model; and   training the classification model based on the sample data to obtain a similarity classification model.   
     
     
         2 . The method according to  claim 1 , wherein the acquiring the user similarity corresponding to each user data pair comprises:
 acquiring biological features of users corresponding to a first user data pair, wherein the first user data pair is any user data pair in the plurality of user data pairs; and   determining a user similarity corresponding to the first user data pair according to the biological features of the users corresponding to the first user data pair.   
     
     
         3 . The method according to  claim 2 , wherein the biological features comprise a facial image feature;
 the acquiring the biological features of the users corresponding to the first user data pair comprises:
 acquiring facial images of the users corresponding to the first user data pair; and 
 performing feature extraction on the facial images to obtain facial image features of the users corresponding to the first user data pair; and 
   the determining the user similarity corresponding to the first user data pair according to the biological features of the users corresponding to the first user data pair comprises:
 determining the user similarity corresponding to the first user data pair according to the facial image features of the users corresponding to the first user data pair. 
   
     
     
         4 . The method according to  claim 2 , wherein the biological features comprise a speech feature;
 the acquiring biological features of users corresponding to the first user data pair comprises:
 acquiring speech data of the users corresponding to the first user data pair; and 
 performing feature extraction on the speech data to obtain speech features of the users corresponding to the first user data pair; and 
   the determining the user similarity corresponding to the first user data pair according to the biological features of the users corresponding to the first user data pair comprises:
 determining the user similarity corresponding to the first user data pair according to the speech features of the users corresponding to the first user data pair. 
   
     
     
         5 . The method according to  claim 1 , wherein the determining, according to the user similarity corresponding to each user data pair and the plurality of user data pairs, the sample data for training the classification model comprises:
 performing feature extraction on each user data pair in the plurality of user data pairs to obtain associated user features between the two sets of user data in each user data pair; and   determining, according to the associated user features between the user data in each user data pair and the user similarity corresponding to each user data pair, the sample data for training the classification model.   
     
     
         6 . The method according to  claim 5 , wherein the determining, according to the associated user features between the two sets of user data in each user data pair and the user similarity corresponding to each user data pair, the sample data for training the classification model comprises:
 selecting positive sample features and negative sample features from user features corresponding to the plurality of user data pairs according to the user similarity corresponding to each user data pair and a predetermined similarity threshold; and   using the positive sample features and the negative sample features as the sample data for training the classification model.   
     
     
         7 . The method according to  claim 6 , wherein the associated user features comprise at least one of a household registration dimension feature, a name dimension feature, a social feature, or an interest feature, wherein
 the household registration dimension feature comprises a feature of user identity information,   the name dimension feature comprises a feature of user name information and a feature of a degree of scarcity of a user surname, and   the social feature comprises a feature of social relationship information of a user.   
     
     
         8 . The method according to  claim 6 , wherein the positive sample features comprise the same quantity of features as the negative sample features. 
     
     
         9 . The method according to  claim 1 , wherein the similarity classification model is a binary classifier model. 
     
     
         10 . A data similarity determining method, comprising:
 acquiring a to-be-detected user data pair, the to-be-detected user data pair including two sets of to-be-detected user data;   performing feature extraction on each set of to-be-detected user data in the to-be-detected user data pair to obtain to-be-detected user features; and   determining a similarity between users corresponding to the two sets of to-be-detected user data in the to-be-detected user data pair according to the to-be-detected user features and a pre-trained similarity classification model.   
     
     
         11 . The method according to  claim 10 , further comprising:
 determining to-be-detected users corresponding to the to-be-detected user data pair as twins if the similarity between the users corresponding to the two sets of to-be-detected user data in the to-be-detected user data pair is greater than a predetermined similarity classification threshold.   
     
     
         12 . A model training device, comprising:
 a processor; and   a memory configured to store instructions,   wherein the processor is configured to execute the instructions to:   acquire a plurality of user data pairs, wherein data fields of two sets of user data in each user data pair have an identical part;   acquire a user similarity corresponding to each user data pair, wherein the user similarity is a similarity between users corresponding to the two sets of user data in each user data pair;   determine, according to the user similarity corresponding to each user data pair and the plurality of user data pairs, sample data for training a preset classification model; and   train the classification model based on the sample data to obtain a similarity classification model.   
     
     
         13 . The device according to  claim 12 , wherein the processor is further configured to execute the instructions to:
 acquire biological features of users corresponding to a first user data pair, wherein the first user data pair is any user data pair in the plurality of user data pairs; and   determine a user similarity corresponding to the first user data pair according to the biological features of the users corresponding to the first user data pair.   
     
     
         14 . The device according to  claim 13 , wherein the biological features comprise a facial image feature, and the processor is further configured to execute the instructions to:
 acquire facial images of the users corresponding to the first user data pair; and   
       perform feature extraction on the facial images to obtain facial image features of the users corresponding to the first user data pair; and
 determine the user similarity corresponding to the first user data pair according to the facial image features of the users corresponding to the first user data pair. 
 
     
     
         15 . The device according to  claim 13 , wherein the biological features comprise a speech feature, and the processor is further configured to execute the instructions to:
 acquire speech data of the users corresponding to the first user data pair; and   
       perform feature extraction on the speech data to obtain speech features of the users corresponding to the first user data pair; and
 determine the user similarity corresponding to the first user data pair according to the speech features of the users corresponding to the first user data pair. 
 
     
     
         16 . The device according to  claim 12 , wherein the processor is further configured to execute the instructions to:
 perform feature extraction on each user data pair in the plurality of user data pairs to obtain associated user features between the two sets of user data in each user data pair; and   determine, according to the associated user features between the two sets of user data in each user data pair and the user similarity corresponding to each user data pair, the sample data for training the classification model.   
     
     
         17 . The device according to  claim 16 , wherein the processor is further configured to execute the instructions to:
 select positive sample features and negative sample features from user features corresponding to the plurality of user data pairs according to the user similarity corresponding to each user data pair and a predetermined similarity threshold; and   use the positive sample features and the negative sample features as the sample data for training the classification model.   
     
     
         18 . The device according to  claim 17 , wherein the associated user features comprise: a household registration dimension feature, a name dimension feature, a social feature, and an interest feature, wherein the household registration dimension feature comprises a feature of user identity information, the name dimension feature comprises a feature of user name information and a feature of a degree of scarcity of a user surname, and the social feature comprises a feature of social relationship information of a user. 
     
     
         19 . The device according to  claim 17 , wherein the positive sample features comprise the same quantity of features as the negative sample features. 
     
     
         20 . The device according to  claim 12 , wherein the similarity classification model is a binary classifier model.

Join the waitlist — get patent alerts

Track US2020012969A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.