Stylized animatable representation
Abstract
A method for computing a stylized, animatable representation of a subject from a family of stylized animatable representations is described. The method comprises accessing a realistic representation of the subject and computing a mesh mapping using a first machine learning model that is trained using a supervised training methodology that uses a set of training examples comprising a plurality of two dimensional (2D) images of the subject and corresponding stylized pictures of the subject. The method also comprises the first trained machine learning model applying the mesh mapping to the realistic representation to produce a target mesh, and selecting the stylized animatable representation from the family, based on closeness of the target mesh with instances of the family.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computerized method of computing a stylized, animatable representation of a subject from a family of stylized animatable representations, the method comprising:
accessing a realistic representation of the subject; computing a mesh mapping using a machine learning model that is trained using a supervised training methodology that uses training examples comprising two dimensional (2D) images of the subject and corresponding stylized pictures of the subject, wherein computing the mesh mapping comprises:
inputting the realistic representation to the trained machine learning model; and
causing the trained machine learning model to compute a mesh mapping;
causing the trained machine learning model to apply the mesh mapping to the realistic representation to produce a target mesh; and selecting the stylized animatable representation from the family, based on closeness of the target mesh with instances of the family.
2 . The method of claim 1 , wherein the trained machine learning model produces an accurate target mesh even when the realistic representation differs from the training examples.
3 . The method of claim 1 , wherein a 2D image is used to create the realistic representation using a technology to reconstruct a 3D model using dense landmarks, another machine learning model being used to predict locations of the dense landmarks in the 2D image, wherein the other trained machine learning model is trained using synthetic training data which gives ground truth landmark annotations.
4 . The method of claim 1 , wherein the training examples are selected as nearest neighbors of the realistic representation.
5 . The method of claim 1 , further comprising, prior to applying the mesh mapping, computing a retopology of the realistic representation.
6 . The method of claim 1 , further comprising computing the mesh mapping by deriving the mesh mapping from a transformation.
7 . The method of claim 1 , wherein the mesh mapping is computed by computing, for the training examples selected as nearest neighbors of the realistic representation, a separate transformation.
8 . An apparatus for computing a stylized, animatable representation of a subject from a family of stylized animatable representations, the apparatus comprising:
a processor; a memory storing a realistic representation of the subject and storing instructions which when executed by the processor cause the processor to: access a realistic representation of the subject; compute a mesh mapping using a machine learning model that is trained using a supervised training methodology that uses training examples comprising two dimensional (2D) images of the subject and corresponding stylized pictures of the subject, wherein computing the mesh mapping comprises:
inputting the realistic representation to the trained machine learning model; and
causing the trained machine learning model to compute a mesh mapping;
cause the trained machine learning model to apply the mesh mapping to the realistic representation to produce a target mesh; and select the stylized animatable representation from the family, based on closeness of the target mesh with instances of the family.
9 . The apparatus of claim 8 , wherein the trained machine learning model produces an accurate target mesh even when the realistic representation differs from the training examples.
10 . The apparatus of claim 8 , wherein a 2D image is used to create the realistic representation using a technology to reconstruct a 3D model using dense landmarks, another trained machine learning model being used to predict locations of the dense landmarks in the 2D image, wherein the other trained machine learning model is trained using synthetic training data which gives ground truth landmark annotations.
11 . The apparatus of claim 8 , wherein the training examples are selected as nearest neighbors of the realistic representation.
12 . The apparatus of claim 8 , wherein the instructions further cause the processor to compute the mesh mapping by deriving the mesh mapping from a transformation.
13 . The apparatus of claim 8 , wherein the instructions further cause the processor to prior to applying the mesh mapping, compute a retopology of the realistic representation.
14 . The apparatus of claim 8 , wherein the mesh mapping is computed by computing, for the training examples selected as nearest neighbors of the realistic representation, a separate transformation.
15 . A non-transitory computer-readable medium embodied with computer-executable instructions that, when executed by a processor, cause the processor to perform operations comprising:
accessing a realistic representation of a subject; computing a mesh mapping using a machine learning model that is trained using a supervised training methodology that uses training examples comprising two dimensional (2D) images of the subject and corresponding stylized pictures of the subject, wherein computing the mesh mapping comprises:
inputting the realistic representation to the trained machine learning model; and
causing the trained machine learning model to compute a mesh mapping;
causing the trained machine learning model to apply the mesh mapping to the realistic representation to produce a target mesh; and selecting a stylized animatable representation from a family of stylized animatable representations, based on closeness of the target mesh with instances of the family.
16 . The non-transitory computer-readable medium of claim 15 , wherein the trained machine learning model produces an accurate target mesh even when the realistic representation differs from the training examples.
17 . The non-transitory computer-readable medium of claim 15 , wherein a 2D image is used to create the realistic representation using a technology to reconstruct a 3D model using dense landmarks, another trained machine learning model being used to predict locations of the dense landmarks in the 2D image, wherein the other trained machine learning model is trained using synthetic training data which gives ground truth landmark annotations.
18 . The non-transitory computer-readable medium of claim 15 , wherein the training examples are selected as nearest neighbors of the realistic representation.
19 . The non-transitory computer-readable medium of claim 15 , wherein the mesh mapping is computed by computing, for the training examples selected as nearest neighbors of the realistic representation, a separate transformation.
20 . The non-transitory computer-readable medium of claim 15 , wherein the computer-executable instructions further cause the processor to perform operations comprising computing the mesh mapping by deriving the mesh mapping from a transformation.Join the waitlist — get patent alerts
Track US2025095262A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.