Image processing
Abstract
A method, apparatus, device, and computer-readable storage medium for image processing are provided. The method includes receiving a text input for an initial image, the text input describing a visual effect for the initial image. A fusion feature for the text input and the initial image is generated based on the text input and the initial image. A target image corresponding to the initial image is generated based on a first image feature of the initial image and the fusion feature, the target image having a visual element related to the visual effect. The fusion of text and image can better express the desired visual effect.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . An image processing method, comprising:
receiving a text input for an initial image, the text input describing a visual effect for the initial image; generating, based on the text input and the initial image, a fusion feature for the text input and the initial image; and generating, based on a first image feature of the initial image and the fusion feature, a target image corresponding to the initial image, the target image having a visual element related to the visual effect.
2 . The method of claim 1 , further comprising:
before generating the target image based on the first image feature and the fusion feature,
generating, based on the text input, a text encoding corresponding to the text input with a text encoder; and
updating the fusion feature with the text encoding.
3 . The method of claim 1 , wherein generating the fusion feature for the text input and the initial image comprises:
determining, based on the text input and the initial image, an initial feature for fusing the text input and the initial image; and determining the fusion feature by converting the initial feature into an initial feature that has a dimension matching the text encoding.
4 . The method of claim 3 , wherein determining the fusion feature by converting the initial feature into the initial feature that has the dimension matching the text encoding comprises:
determining, based on the initial feature, a key feature and a value feature for an attention mechanism; and determining, based on the key feature, the value feature, and a predetermined query feature, the fusion feature with the attention mechanism.
5 . The method of claim 3 , wherein determining the initial feature for fusing the text input and the initial image comprises:
providing the text input and the initial image as input to a multimodal model to obtain an output of a predetermined intermediate layer of the multimodal model; and determining the initial feature based on the output of the predetermined intermediate layer.
6 . The method of claim 1 , wherein generating the target image corresponding to the initial image comprises:
generating, based on the initial image, a second image feature of the initial image with a control model; and generating the target image based on the first image feature, the second image feature, and the fusion feature.
7 . The method of claim 6 , wherein generating the target image based on the first image feature, the second image feature, and the fusion feature comprises:
generating, based on the first image feature, the target image by using the second image feature and the fusion feature as a control condition.
8 . The method of claim 1 , wherein the first image feature of the initial image is determined by:
generating, based on the initial image and a noise signal, the first image feature with an image encoder.
9 . An electronic device, comprising:
at least one processor; and at least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform acts comprising:
receiving a text input for an initial image, the text input describing a visual effect for the initial image;
generating, based on the text input and the initial image, a fusion feature for the text input and the initial image; and
generating, based on a first image feature of the initial image and the fusion feature, a target image corresponding to the initial image, the target image having a visual element related to the visual effect.
10 . The electronic device of claim 9 , wherein the acts further comprise:
before generating the target image based on the first image feature and the fusion feature, generating, based on the text input, a text encoding corresponding to the text input with a text encoder; and updating the fusion feature with the text encoding.
11 . The electronic device of claim 9 , wherein generating the fusion feature for the text input and the initial image comprises:
determining, based on the text input and the initial image, an initial feature for fusing the text input and the initial image; and determining the fusion feature by converting the initial feature into an initial feature that has a dimension matching the text encoding.
12 . The electronic device of claim 11 , wherein determining the fusion feature by converting the initial feature into the initial feature that has the dimension matching the text encoding comprises:
determining, based on the initial feature, a key feature and a value feature for an attention mechanism; and determining, based on the key feature, the value feature, and a predetermined query feature, the fusion feature with the attention mechanism.
13 . The electronic device of claim 11 , wherein determining the initial feature for fusing the text input and the initial image comprises:
providing the text input and the initial image as input to a multimodal model to obtain an output of a predetermined intermediate layer of the multimodal model; and determining the initial feature based on the output of the predetermined intermediate layer.
14 . The electronic device of claim 9 , wherein generating the target image corresponding to the initial image comprises:
generating, based on the initial image, a second image feature of the initial image with a control model; and generating the target image based on the first image feature, the second image feature, and the fusion feature.
15 . The electronic device of claim 14 , wherein generating the target image based on the first image feature, the second image feature, and the fusion feature comprises:
generating, based on the first image feature, the target image by using the second image feature and the fusion feature as a control condition.
16 . The electronic device of claim 9 , wherein the first image feature of the initial image is determined by:
generating, based on the initial image and a noise signal, the first image feature with an image encoder.
17 . A non-transitory computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to perform acts comprising:
receiving a text input for an initial image, the text input describing a visual effect for the initial image; generating, based on the text input and the initial image, a fusion feature for the text input and the initial image; and generating, based on a first image feature of the initial image and the fusion feature, a target image corresponding to the initial image, the target image having a visual element related to the visual effect.
18 . The non-transitory computer-readable storage medium of claim 17 , wherein the acts further comprise:
before generating the target image based on the first image feature and the fusion feature, generating, based on the text input, a text encoding corresponding to the text input with a text encoder; and updating the fusion feature with the text encoding.
19 . The non-transitory computer-readable storage medium of claim 17 , wherein generating the fusion feature for the text input and the initial image comprises:
determining, based on the text input and the initial image, an initial feature for fusing the text input and the initial image; and determining the fusion feature by converting the initial feature into an initial feature that has a dimension matching the text encoding.
20 . The non-transitory computer-readable storage medium of claim 19 , wherein determining the fusion feature by converting the initial feature into the initial feature that has the dimension matching the text encoding comprises:
determining, based on the initial feature, a key feature and a value feature for an attention mechanism; and determining, based on the key feature, the value feature, and a predetermined query feature, the fusion feature with the attention mechanism.Join the waitlist — get patent alerts
Track US2026065525A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.