Method for Training Image Generation Model, Method for Generating Digital Human Image, Electronic Device and Storage Medium
Abstract
A method for training an image generation model, a method for generating a digital human image, and related apparatuses are provided, relating to the fields of artificial intelligence, big model, big data and other technologies. The method for training an image generation model includes: obtaining N target facial images of a target face, wherein N is an integer greater than 1; inputting the N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after the target face is fused with each target background image; and training the preset image generation model based on a degree of difference between a first facial feature in the target digital human image and a second facial feature of the target face in the target facial images to obtain a target image generation model.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for training an image generation model, comprising:
obtaining N target facial images of a target face, wherein N is an integer greater than 1; inputting the N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after the target face is fused with each target background image; and training the preset image generation model based on a degree of difference between a first facial feature in the target digital human image and a second facial feature of the target face in the target facial images to obtain a target image generation model.
2 . The method of claim 1 , wherein obtaining the N target facial images of the target face, comprises:
obtaining M initial facial images of the target face, wherein M is a natural number less than or equal to N; and obtaining at least one facial expansion image for the target face in at least one of following ways, to obtain the N target facial images based on the facial expansion image: perturbing the target face locally based on facial key features of the initial facial images; fine-tuning a viewing angle of the target face based on facial key features of the initial facial images; or transforming light environment of the target face based on facial key features of the initial facial images.
3 . The method of claim 2 , wherein perturbing the target face locally based on the facial key features of the initial facial images, comprises:
fine-tuning non-key points in the target face for local perturbation based on the facial key features of the initial facial images.
4 . The method of claim 2 , wherein fine-tuning the viewing angle of the target face based on the facial key features of the initial facial images, comprises:
fine-tuning the facial key features of the initial facial images based on the target face at a preset viewing angle, to fine-tune the viewing angle of the target face.
5 . The method of claim 2 , wherein transforming the light environment of the target face based on the facial key features of the initial facial images, comprises:
fine-tuning the facial key features of the initial facial images based on a preset light condition, to transform the light environment of the target face.
6 . The method of claim 2 , wherein different target facial images have different image features.
7 . The method of claim 2 , further comprising:
preprocessing the initial facial images; performing feature encoding on preprocessed initial facial images; and performing feature extraction on key points in feature-encoded initial facial images to obtain the facial key features of the initial facial images.
8 . The method of claim 2 , wherein inputting the N target facial images and the at least one target background image into the preset image generation model to obtain the target digital human image after the target face is fused with each target background image, comprises:
inputting the N target facial images and at least one target background image into an image generation network of the preset image generation model, to extract facial features from the N target facial images to obtain a facial feature set for the N target facial images, and extract background features from each target background image to obtain background features for each target background image; and inputting the facial feature set and the background features of each target background image into a consistency fusion network of the preset image generation model, to perform facial consistency constraint on the facial feature set of the N target facial images, and perform feature fusion with the background features of each target background image after the facial consistency constraint to obtain the target digital human image after the target face is fused with each target background image.
9 . The method of claim 8 , wherein training the preset image generation model based on the degree of difference between the first facial feature in the target digital human image and the second facial feature of the target face in the target facial images, comprises:
calculating at least similarity between the first facial feature in the target digital human image and the second facial feature of the target face in each target facial image to obtain similarity information; and fine-tuning an image generation parameter in the preset image generation model based on the similarity information.
10 . A method for generating a digital human image, comprising:
obtaining a plurality of facial images to be processed for a preset face; and inputting the plurality of facial images to be processed and at least one preset background image into a target image generation model to obtain a digital human image after the preset face is fused with each preset background image; wherein the target image generation model is obtained by training a preset image generation model based on a degree of difference between a first facial feature in a target digital human image and a second facial feature of a target face in a target facial image; the target digital human image is obtained by inputting at least N target facial images into the preset image generation model; the N target facial images are obtained by expanding M initial facial images; N is an integer greater than 1; and M is a natural number less than or equal to N.
11 . An electronic device, comprising:
at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores an instruction executable by the at least one processor, and the instruction, when executed by the at least one processor, enables the at least one processor to execute: obtaining N target facial images of a target face, wherein N is an integer greater than 1; inputting the N target facial images and at least one target background image into a preset image generation model to obtain a target digital human image after the target face is fused with each target background image; and training the preset image generation model based on a degree of difference between a first facial feature in the target digital human image and a second facial feature of the target face in the target facial images to obtain a target image generation model.
12 . The electronic device of claim 11 , wherein the instruction, when executed by the at least one processor, enables the at least one processor to execute obtaining the N target facial images of the target face, by:
obtaining M initial facial images of the target face, wherein M is a natural number less than or equal to N; and obtaining at least one facial expansion image for the target face in at least one of following ways, to obtain the N target facial images based on the facial expansion image: perturbing the target face locally based on facial key features of the initial facial images; fine-tuning a viewing angle of the target face based on facial key features of the initial facial images; or transforming light environment of the target face based on facial key features of the initial facial images.
13 . The electronic device of claim 12 , wherein the instruction, when executed by the at least one processor, enables the at least one processor to execute perturbing the target face locally based on the facial key features of the initial facial images, by:
fine-tuning non-key points in the target face for local perturbation based on the facial key features of the initial facial images.
14 . The electronic device of claim 12 , wherein the instruction, when executed by the at least one processor, enables the at least one processor to execute fine-tuning the viewing angle of the target face based on the facial key features of the initial facial images, by:
fine-tuning the facial key features of the initial facial images based on the target face at a preset viewing angle, to fine-tune the viewing angle of the target face.
15 . The electronic device of claim 12 , wherein the instruction, when executed by the at least one processor, enables the at least one processor to execute transforming the light environment of the target face based on the facial key features of the initial facial images, by:
fine-tuning the facial key features of the initial facial images based on a preset light condition, to transform the light environment of the target face.
16 . The electronic device of claim 12 , wherein different target facial images have different image features.
17 . The electronic device of claim 12 , wherein the instruction, when executed by the at least one processor, enables the at least one processor to further execute:
preprocessing the initial facial images; performing feature encoding on preprocessed initial facial images; and performing feature extraction on key points in feature-encoded initial facial images to obtain the facial key features of the initial facial images.
18 . An electronic device, comprising:
at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores an instruction executable by the at least one processor, and the instruction, when executed by the at least one processor, enables the at least one processor to execute the method of claim 10 .
19 . A non-transitory computer-readable storage medium storing a computer instruction thereon, wherein the computer instruction is used to cause a computer to execute the method of claim 1 .
20 . A non-transitory computer-readable storage medium storing a computer instruction thereon, wherein the computer instruction is used to cause a computer to execute the method of claim 10 .Join the waitlist — get patent alerts
Track US2025315989A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.