Neural network target feature detection
Abstract
A method of training a neural network for detecting target features in images is described. The neural network is trained using a first data set that includes labeled images, where at least some of the labeled images having subjects with labeled features, including: dividing each of the labeled images of the first data set into a respective plurality of tiles, and generating, for each of the plurality of tiles, a plurality of feature anchors that indicate target features within the corresponding tile. Target features that correspond to the plurality of feature anchors are detected in a second data set of unlabeled images. Images of the second data set having target features that were not detected are labeled. A third data set that includes the first data set and the labeled images of the second data set is generated. The neural network is trained using the third data set.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method of training a neural network for detecting target features in images, the method comprising:
training the neural network using a first data set that includes labeled images, at least some of the labeled images having subjects with labeled features, including
dividing each of the labeled images of the first data set into a respective plurality of tiles, and
generating, for each of the plurality of tiles, a plurality of feature anchors that indicate target features within the corresponding tile;
detecting target features that correspond to the plurality of feature anchors in a second data set of unlabeled images; labeling images of the second data set having target features that were not detected; generating a third data set that includes the first data set and the labeled images of the second data set; training the neural network using the third data set.
2 . The computer-implemented method of claim 1 , wherein each of the plurality of feature anchors indicates a bounding box within the corresponding tile that contains a target feature.
3 . The computer-implemented method of claim 1 , further comprising generating the third data set to include images of the second data set that correspond to false positive detections of the target features.
4 . The computer-implemented method of claim 3 , further comprising generating the second data set to include images from videos without people.
5 . The computer-implemented method of claim 1 , further comprising generating the second data set to include images from videos that depict subjects with different head poses.
6 . The computer-implemented method of claim 1 , wherein generating the third data set comprises performing a randomized crop of different aspect ratios on at least some of the third data set.
7 . The computer-implemented method of claim 1 , wherein generating the third data set comprises generating at least some images having light levels below a low light threshold, including augmenting an image of the third data set to have light levels below the low light threshold.
8 . The computer-implemented method of claim 1 , wherein generating the third data set comprises generating at least some images having partially occluded target features, including augmenting an image of the third data set to have a partially occluded target feature.
9 . The computer-implemented method of claim 8 , wherein augmenting the image comprises cropping the image or inserting a block to obtain the partially occluded target feature.
10 . The computer-implemented method of claim 1 , wherein training the neural network using the third data set comprises training the neural network using floating point values for weights of the neural network;
the method further comprises quantizing the weights of the neural network using integers.
11 . The computer-implemented method of claim 1 , wherein training the neural network using the first data set of labeled images comprises normalizing RGB values of the labeled images from 0 to 255 to −1 to 1.
12 . The computer-implemented method of claim 1 , wherein the target features include a subject face and/or subject body.
13 . A system for training a neural network for detecting target features in images, the system comprising:
a processor, and a memory storing computer-executable instructions that when executed by the processor cause the system to: train the neural network using a first data set that includes labeled images, at least some of the labeled images having subjects with labeled features, including
dividing each of the labeled images into a respective plurality of tiles, and
generating, for each of the plurality of tiles, a plurality of feature anchors that indicate target features within the corresponding tile;
detecting target features that correspond to the plurality of feature anchors in a second data set of unlabeled images; labeling images of the second data set having target features that were not detected; generating a third data set that includes the first data set of labeled images and the labeled images of the second data set; and training the neural network using the third data set.
14 . The system of claim 13 , wherein each of the plurality of feature anchors indicates a bounding box within the corresponding tile that contains a target feature.
15 . The system of claim 13 , further comprising generating the third data set to include images of the second data set that correspond to false positive detections of the target features.
16 . The system of claim 15 , further comprising generating the second data set to include images from videos without people.
17 . The system of claim 13 , further comprising generating the second data set to include images from videos that depict subjects with different head poses.
18 . The system of claim 13 , wherein generating the third data set comprises performing a randomized crop of different aspect ratios on at least some of the third data set.
19 . An image processing system that includes a neural network implemented on a computer for feature detection, comprising:
a convolutional neural network having a plurality of layers stacked sequentially, including:
a first set of layers, each layer of the first set of layers having a depth-wise convolution and a point-wise convolution, wherein the first set of layers is a first subset of a different neural network; and
a second set of layers after the first set of layers, each layer of the second set of layers having a point-wise convolution.
20 . The image processing system of claim 19 , wherein the second set of layers is based upon a compression of a remainder subset of the different neural network.Join the waitlist — get patent alerts
Track US2022358332A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.