US2022358332A1PendingUtilityA1

Neural network target feature detection

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: May 7, 2021Filed: May 7, 2021Published: Nov 10, 2022
Est. expiryMay 7, 2041(~14.8 yrs left)· nominal 20-yr term from priority
G06N 3/045G06F 18/2148G06F 18/2163G06F 18/211G06F 18/2155G06N 3/08G06V 2201/07G06V 40/161G06V 40/10G06N 3/04G06K 9/6228G06K 9/00362G06K 9/6261G06K 2209/21G06K 9/6257G06K 9/00228G06K 9/6259G06N 3/0464G06N 3/09G06N 3/0895G06N 3/0495G06V 10/255G06V 10/454G06V 10/82
55
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method of training a neural network for detecting target features in images is described. The neural network is trained using a first data set that includes labeled images, where at least some of the labeled images having subjects with labeled features, including: dividing each of the labeled images of the first data set into a respective plurality of tiles, and generating, for each of the plurality of tiles, a plurality of feature anchors that indicate target features within the corresponding tile. Target features that correspond to the plurality of feature anchors are detected in a second data set of unlabeled images. Images of the second data set having target features that were not detected are labeled. A third data set that includes the first data set and the labeled images of the second data set is generated. The neural network is trained using the third data set.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method of training a neural network for detecting target features in images, the method comprising:
 training the neural network using a first data set that includes labeled images, at least some of the labeled images having subjects with labeled features, including
 dividing each of the labeled images of the first data set into a respective plurality of tiles, and 
 generating, for each of the plurality of tiles, a plurality of feature anchors that indicate target features within the corresponding tile; 
   detecting target features that correspond to the plurality of feature anchors in a second data set of unlabeled images;   labeling images of the second data set having target features that were not detected;   generating a third data set that includes the first data set and the labeled images of the second data set;   training the neural network using the third data set.   
     
     
         2 . The computer-implemented method of  claim 1 , wherein each of the plurality of feature anchors indicates a bounding box within the corresponding tile that contains a target feature. 
     
     
         3 . The computer-implemented method of  claim 1 , further comprising generating the third data set to include images of the second data set that correspond to false positive detections of the target features. 
     
     
         4 . The computer-implemented method of  claim 3 , further comprising generating the second data set to include images from videos without people. 
     
     
         5 . The computer-implemented method of  claim 1 , further comprising generating the second data set to include images from videos that depict subjects with different head poses. 
     
     
         6 . The computer-implemented method of  claim 1 , wherein generating the third data set comprises performing a randomized crop of different aspect ratios on at least some of the third data set. 
     
     
         7 . The computer-implemented method of  claim 1 , wherein generating the third data set comprises generating at least some images having light levels below a low light threshold, including augmenting an image of the third data set to have light levels below the low light threshold. 
     
     
         8 . The computer-implemented method of  claim 1 , wherein generating the third data set comprises generating at least some images having partially occluded target features, including augmenting an image of the third data set to have a partially occluded target feature. 
     
     
         9 . The computer-implemented method of  claim 8 , wherein augmenting the image comprises cropping the image or inserting a block to obtain the partially occluded target feature. 
     
     
         10 . The computer-implemented method of  claim 1 , wherein training the neural network using the third data set comprises training the neural network using floating point values for weights of the neural network;
 the method further comprises quantizing the weights of the neural network using integers.   
     
     
         11 . The computer-implemented method of  claim 1 , wherein training the neural network using the first data set of labeled images comprises normalizing RGB values of the labeled images from 0 to 255 to −1 to 1. 
     
     
         12 . The computer-implemented method of  claim 1 , wherein the target features include a subject face and/or subject body. 
     
     
         13 . A system for training a neural network for detecting target features in images, the system comprising:
 a processor, and   a memory storing computer-executable instructions that when executed by the processor cause the system to:   train the neural network using a first data set that includes labeled images, at least some of the labeled images having subjects with labeled features, including
 dividing each of the labeled images into a respective plurality of tiles, and 
 generating, for each of the plurality of tiles, a plurality of feature anchors that indicate target features within the corresponding tile; 
   detecting target features that correspond to the plurality of feature anchors in a second data set of unlabeled images;   labeling images of the second data set having target features that were not detected;   generating a third data set that includes the first data set of labeled images and the labeled images of the second data set; and   training the neural network using the third data set.   
     
     
         14 . The system of  claim 13 , wherein each of the plurality of feature anchors indicates a bounding box within the corresponding tile that contains a target feature. 
     
     
         15 . The system of  claim 13 , further comprising generating the third data set to include images of the second data set that correspond to false positive detections of the target features. 
     
     
         16 . The system of  claim 15 , further comprising generating the second data set to include images from videos without people. 
     
     
         17 . The system of  claim 13 , further comprising generating the second data set to include images from videos that depict subjects with different head poses. 
     
     
         18 . The system of  claim 13 , wherein generating the third data set comprises performing a randomized crop of different aspect ratios on at least some of the third data set. 
     
     
         19 . An image processing system that includes a neural network implemented on a computer for feature detection, comprising:
 a convolutional neural network having a plurality of layers stacked sequentially, including:
 a first set of layers, each layer of the first set of layers having a depth-wise convolution and a point-wise convolution, wherein the first set of layers is a first subset of a different neural network; and 
 a second set of layers after the first set of layers, each layer of the second set of layers having a point-wise convolution. 
   
     
     
         20 . The image processing system of  claim 19 , wherein the second set of layers is based upon a compression of a remainder subset of the different neural network.

Join the waitlist — get patent alerts

Track US2022358332A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.