US2025095262A1PendingUtilityA1

Stylized animatable representation

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Nov 22, 2022Filed: Nov 26, 2024Published: Mar 20, 2025
Est. expiryNov 22, 2042(~16.3 yrs left)· nominal 20-yr term from priority
G06T 17/20G06V 10/761G06V 10/762G06V 20/653G06T 13/40G06T 17/00
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method for computing a stylized, animatable representation of a subject from a family of stylized animatable representations is described. The method comprises accessing a realistic representation of the subject and computing a mesh mapping using a first machine learning model that is trained using a supervised training methodology that uses a set of training examples comprising a plurality of two dimensional (2D) images of the subject and corresponding stylized pictures of the subject. The method also comprises the first trained machine learning model applying the mesh mapping to the realistic representation to produce a target mesh, and selecting the stylized animatable representation from the family, based on closeness of the target mesh with instances of the family.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computerized method of computing a stylized, animatable representation of a subject from a family of stylized animatable representations, the method comprising:
 accessing a realistic representation of the subject;   computing a mesh mapping using a machine learning model that is trained using a supervised training methodology that uses training examples comprising two dimensional (2D) images of the subject and corresponding stylized pictures of the subject, wherein computing the mesh mapping comprises:
 inputting the realistic representation to the trained machine learning model; and 
 causing the trained machine learning model to compute a mesh mapping; 
   causing the trained machine learning model to apply the mesh mapping to the realistic representation to produce a target mesh; and   selecting the stylized animatable representation from the family, based on closeness of the target mesh with instances of the family.   
     
     
         2 . The method of  claim 1 , wherein the trained machine learning model produces an accurate target mesh even when the realistic representation differs from the training examples. 
     
     
         3 . The method of  claim 1 , wherein a 2D image is used to create the realistic representation using a technology to reconstruct a 3D model using dense landmarks, another machine learning model being used to predict locations of the dense landmarks in the 2D image, wherein the other trained machine learning model is trained using synthetic training data which gives ground truth landmark annotations. 
     
     
         4 . The method of  claim 1 , wherein the training examples are selected as nearest neighbors of the realistic representation. 
     
     
         5 . The method of  claim 1 , further comprising, prior to applying the mesh mapping, computing a retopology of the realistic representation. 
     
     
         6 . The method of  claim 1 , further comprising computing the mesh mapping by deriving the mesh mapping from a transformation. 
     
     
         7 . The method of  claim 1 , wherein the mesh mapping is computed by computing, for the training examples selected as nearest neighbors of the realistic representation, a separate transformation. 
     
     
         8 . An apparatus for computing a stylized, animatable representation of a subject from a family of stylized animatable representations, the apparatus comprising:
 a processor;   a memory storing a realistic representation of the subject and storing instructions which when executed by the processor cause the processor to:   access a realistic representation of the subject;   compute a mesh mapping using a machine learning model that is trained using a supervised training methodology that uses training examples comprising two dimensional (2D) images of the subject and corresponding stylized pictures of the subject, wherein computing the mesh mapping comprises:
 inputting the realistic representation to the trained machine learning model; and 
 causing the trained machine learning model to compute a mesh mapping; 
   cause the trained machine learning model to apply the mesh mapping to the realistic representation to produce a target mesh; and   select the stylized animatable representation from the family, based on closeness of the target mesh with instances of the family.   
     
     
         9 . The apparatus of  claim 8 , wherein the trained machine learning model produces an accurate target mesh even when the realistic representation differs from the training examples. 
     
     
         10 . The apparatus of  claim 8 , wherein a 2D image is used to create the realistic representation using a technology to reconstruct a 3D model using dense landmarks, another trained machine learning model being used to predict locations of the dense landmarks in the 2D image, wherein the other trained machine learning model is trained using synthetic training data which gives ground truth landmark annotations. 
     
     
         11 . The apparatus of  claim 8 , wherein the training examples are selected as nearest neighbors of the realistic representation. 
     
     
         12 . The apparatus of  claim 8 , wherein the instructions further cause the processor to compute the mesh mapping by deriving the mesh mapping from a transformation. 
     
     
         13 . The apparatus of  claim 8 , wherein the instructions further cause the processor to prior to applying the mesh mapping, compute a retopology of the realistic representation. 
     
     
         14 . The apparatus of  claim 8 , wherein the mesh mapping is computed by computing, for the training examples selected as nearest neighbors of the realistic representation, a separate transformation. 
     
     
         15 . A non-transitory computer-readable medium embodied with computer-executable instructions that, when executed by a processor, cause the processor to perform operations comprising:
 accessing a realistic representation of a subject;   computing a mesh mapping using a machine learning model that is trained using a supervised training methodology that uses training examples comprising two dimensional (2D) images of the subject and corresponding stylized pictures of the subject, wherein computing the mesh mapping comprises:
 inputting the realistic representation to the trained machine learning model; and 
 causing the trained machine learning model to compute a mesh mapping; 
   causing the trained machine learning model to apply the mesh mapping to the realistic representation to produce a target mesh; and   selecting a stylized animatable representation from a family of stylized animatable representations, based on closeness of the target mesh with instances of the family.   
     
     
         16 . The non-transitory computer-readable medium of  claim 15 , wherein the trained machine learning model produces an accurate target mesh even when the realistic representation differs from the training examples. 
     
     
         17 . The non-transitory computer-readable medium of  claim 15 , wherein a 2D image is used to create the realistic representation using a technology to reconstruct a 3D model using dense landmarks, another trained machine learning model being used to predict locations of the dense landmarks in the 2D image, wherein the other trained machine learning model is trained using synthetic training data which gives ground truth landmark annotations. 
     
     
         18 . The non-transitory computer-readable medium of  claim 15 , wherein the training examples are selected as nearest neighbors of the realistic representation. 
     
     
         19 . The non-transitory computer-readable medium of  claim 15 , wherein the mesh mapping is computed by computing, for the training examples selected as nearest neighbors of the realistic representation, a separate transformation. 
     
     
         20 . The non-transitory computer-readable medium of  claim 15 , wherein the computer-executable instructions further cause the processor to perform operations comprising computing the mesh mapping by deriving the mesh mapping from a transformation.

Join the waitlist — get patent alerts

Track US2025095262A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.