US2024212325A1PendingUtilityA1

Systems and Methods for Training Models to Predict Dense Correspondences in Images Using Geodesic Distances

Assignee: GOOGLE LLCPriority: Mar 11, 2021Filed: Mar 6, 2024Published: Jun 27, 2024
Est. expiryMar 11, 2041(~14.6 yrs left)· nominal 20-yr term from priority
G06N 3/0464G06N 3/0895G06V 20/653G06V 10/82G06F 18/22G06F 18/2413G06F 18/214G06T 2207/20084G06T 2207/20081G06T 17/00G06V 10/44G06V 10/751G06T 7/70G06T 2207/30196G06T 2207/30124G06N 3/084G06T 17/20G06V 10/771G06T 7/33
70
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for training models to predict dense correspondences across images such as human images. A model may be trained using synthetic training data created from one or more 3D computer models of a subject. In addition, one or more geodesic distances derived from the surfaces of one or more of the 3D models may be used to generate one or more loss values, which may in turn be used in modifying the model's parameters during training.

Claims

exact text as granted — not AI-modified
1 . A method of training a neural network, the method comprising:
 determining, by one or more processors, a first feature distance between a first point as represented in a first feature map and a second point as represented in a second feature map, the first point and the second point corresponding to the same feature on a three-dimensional model of a subject, the first feature map being based on a first image of the subject, and the second feature map being based on a second image of the subject;   determining, by the one or more processors, a first geodesic distance between a first pair of selected points as represented in a surface map corresponding to the first image;   determining, by the one or more processors, a second geodesic distance between a second pair of selected points as represented in the surface map; and   modifying, by the one or more processors, one or more parameters of the neural network based at least in part on a pair of loss values, a first one of the pair of loss values being based on the first feature distance, and a second one of the pair of loss values being based on at least the first and second geodesic distances.   
     
     
         2 . The method of  claim 1 , wherein the first loss value is further based on a set of additional feature distances. 
     
     
         3 . The method of  claim 2 , wherein each given feature distance of the set of additional feature distances is between a selected point as represented in the first feature map and a corresponding point as represented in the second feature map, the selected point and the corresponding point corresponding to the same feature on the three-dimensional model of the subject. 
     
     
         4 . The method of  claim 2 , wherein the first point and each selected point collectively represent all pixels in the first image. 
     
     
         5 . The method of  claim 1 , wherein the second loss value is further based on at least one additional pair of feature distances and at least one additional pair of geodesic distances. 
     
     
         6 . The method of  claim 1 , further comprising generating, via the neural network, the first feature map. 
     
     
         7 . The method of  claim 6 , further comprising generating the second feature map. 
     
     
         8 . The method of  claim 7 , wherein generating the first feature map and generating the second feature map are performed using the three-dimensional model of the subject. 
     
     
         9 . The method of  claim 1 , wherein the first point, when represented in a second surface map, corresponds to a feature on the three-dimensional model of the subject that is not represented in the second feature map. 
     
     
         10 . The method of  claim 1 , the method further comprising the one or more processors generating one of the first image or the second image. 
     
     
         11 . The method of  claim 10 , the method further comprising the one or more processors generating the other one of the first image or the second image. 
     
     
         12 . The method of  claim 1 , the method further comprising the one or more processors generating the first surface map. 
     
     
         13 . The method of  claim 1 , wherein the subject is a human or a representation of a human. 
     
     
         14 . The method of  claim 1 , wherein the subject is in a different pose in the first image than in the second image. 
     
     
         15 . The method of  claim 1 , wherein the first image is generated from a different perspective of the three-dimensional model of the subject than the second image. 
     
     
         16 . A processing system comprising:
 memory storing a neural network; and   one or more processors operatively coupled to the memory and configured to use the neural network to predict correspondences in images, wherein the neural network has been trained to predict correspondences in images pursuant to a training method comprising:
 determining a first feature distance between a first point as represented in a first feature map and a second point as represented in a second feature map, the first point and the second point corresponding to the same feature on a three-dimensional model of a subject, the first feature map being based on a first image of the subject, and the second feature map being based on a second image of the subject; 
 determining a first geodesic distance between a first pair of selected points as represented in a surface map corresponding to the first image; 
 determining a second geodesic distance between a second pair of selected points as represented in the surface map; and 
 modifying one or more parameters of the neural network based at least in part on a pair of loss values, a first one of the pair of loss values being based on the first feature distance, and a second one of the pair of loss values being based on at least the first and second geodesic distances. 
   
     
     
         17 . The processing system of  claim 16 , wherein the one or more processor are further configured to generate, via the neural network, the first feature map. 
     
     
         18 . The processing system of  claim 16 , wherein the one or more processors are further configured to generate at least one of the first image or the second image. 
     
     
         19 . The processing system of  claim 18 , wherein the one or more processors are further configured to generate the first surface map. 
     
     
         20 . The processing system of  claim 16 , wherein the first image is generated from a different perspective of the three-dimensional model of the subject than the second image.

Join the waitlist — get patent alerts

Track US2024212325A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.