US2026065525A1PendingUtilityA1

Image processing

Assignee: BEIJING ZITIAO NETWORK TECHNOLOGY CO LTDPriority: Aug 30, 2024Filed: Aug 20, 2025Published: Mar 5, 2026
Est. expiryAug 30, 2044(~18.1 yrs left)· nominal 20-yr term from priority
G06T 11/60G06T 11/00
67
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method, apparatus, device, and computer-readable storage medium for image processing are provided. The method includes receiving a text input for an initial image, the text input describing a visual effect for the initial image. A fusion feature for the text input and the initial image is generated based on the text input and the initial image. A target image corresponding to the initial image is generated based on a first image feature of the initial image and the fusion feature, the target image having a visual element related to the visual effect. The fusion of text and image can better express the desired visual effect.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . An image processing method, comprising:
 receiving a text input for an initial image, the text input describing a visual effect for the initial image;   generating, based on the text input and the initial image, a fusion feature for the text input and the initial image; and   generating, based on a first image feature of the initial image and the fusion feature, a target image corresponding to the initial image, the target image having a visual element related to the visual effect.   
     
     
         2 . The method of  claim 1 , further comprising:
 before generating the target image based on the first image feature and the fusion feature,
 generating, based on the text input, a text encoding corresponding to the text input with a text encoder; and 
 updating the fusion feature with the text encoding. 
   
     
     
         3 . The method of  claim 1 , wherein generating the fusion feature for the text input and the initial image comprises:
 determining, based on the text input and the initial image, an initial feature for fusing the text input and the initial image; and   determining the fusion feature by converting the initial feature into an initial feature that has a dimension matching the text encoding.   
     
     
         4 . The method of  claim 3 , wherein determining the fusion feature by converting the initial feature into the initial feature that has the dimension matching the text encoding comprises:
 determining, based on the initial feature, a key feature and a value feature for an attention mechanism; and   determining, based on the key feature, the value feature, and a predetermined query feature, the fusion feature with the attention mechanism.   
     
     
         5 . The method of  claim 3 , wherein determining the initial feature for fusing the text input and the initial image comprises:
 providing the text input and the initial image as input to a multimodal model to obtain an output of a predetermined intermediate layer of the multimodal model; and   determining the initial feature based on the output of the predetermined intermediate layer.   
     
     
         6 . The method of  claim 1 , wherein generating the target image corresponding to the initial image comprises:
 generating, based on the initial image, a second image feature of the initial image with a control model; and   generating the target image based on the first image feature, the second image feature, and the fusion feature.   
     
     
         7 . The method of  claim 6 , wherein generating the target image based on the first image feature, the second image feature, and the fusion feature comprises:
 generating, based on the first image feature, the target image by using the second image feature and the fusion feature as a control condition.   
     
     
         8 . The method of  claim 1 , wherein the first image feature of the initial image is determined by:
 generating, based on the initial image and a noise signal, the first image feature with an image encoder.   
     
     
         9 . An electronic device, comprising:
 at least one processor; and   at least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform acts comprising:
 receiving a text input for an initial image, the text input describing a visual effect for the initial image; 
 generating, based on the text input and the initial image, a fusion feature for the text input and the initial image; and 
 generating, based on a first image feature of the initial image and the fusion feature, a target image corresponding to the initial image, the target image having a visual element related to the visual effect. 
   
     
     
         10 . The electronic device of  claim 9 , wherein the acts further comprise:
 before generating the target image based on the first image feature and the fusion feature,   generating, based on the text input, a text encoding corresponding to the text input with a text encoder; and   updating the fusion feature with the text encoding.   
     
     
         11 . The electronic device of  claim 9 , wherein generating the fusion feature for the text input and the initial image comprises:
 determining, based on the text input and the initial image, an initial feature for fusing the text input and the initial image; and   determining the fusion feature by converting the initial feature into an initial feature that has a dimension matching the text encoding.   
     
     
         12 . The electronic device of  claim 11 , wherein determining the fusion feature by converting the initial feature into the initial feature that has the dimension matching the text encoding comprises:
 determining, based on the initial feature, a key feature and a value feature for an attention mechanism; and   determining, based on the key feature, the value feature, and a predetermined query feature, the fusion feature with the attention mechanism.   
     
     
         13 . The electronic device of  claim 11 , wherein determining the initial feature for fusing the text input and the initial image comprises:
 providing the text input and the initial image as input to a multimodal model to obtain an output of a predetermined intermediate layer of the multimodal model; and   determining the initial feature based on the output of the predetermined intermediate layer.   
     
     
         14 . The electronic device of  claim 9 , wherein generating the target image corresponding to the initial image comprises:
 generating, based on the initial image, a second image feature of the initial image with a control model; and   generating the target image based on the first image feature, the second image feature, and the fusion feature.   
     
     
         15 . The electronic device of  claim 14 , wherein generating the target image based on the first image feature, the second image feature, and the fusion feature comprises:
 generating, based on the first image feature, the target image by using the second image feature and the fusion feature as a control condition.   
     
     
         16 . The electronic device of  claim 9 , wherein the first image feature of the initial image is determined by:
 generating, based on the initial image and a noise signal, the first image feature with an image encoder.   
     
     
         17 . A non-transitory computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to perform acts comprising:
 receiving a text input for an initial image, the text input describing a visual effect for the initial image;   generating, based on the text input and the initial image, a fusion feature for the text input and the initial image; and   generating, based on a first image feature of the initial image and the fusion feature, a target image corresponding to the initial image, the target image having a visual element related to the visual effect.   
     
     
         18 . The non-transitory computer-readable storage medium of  claim 17 , wherein the acts further comprise:
 before generating the target image based on the first image feature and the fusion feature,   generating, based on the text input, a text encoding corresponding to the text input with a text encoder; and   updating the fusion feature with the text encoding.   
     
     
         19 . The non-transitory computer-readable storage medium of  claim 17 , wherein generating the fusion feature for the text input and the initial image comprises:
 determining, based on the text input and the initial image, an initial feature for fusing the text input and the initial image; and   determining the fusion feature by converting the initial feature into an initial feature that has a dimension matching the text encoding.   
     
     
         20 . The non-transitory computer-readable storage medium of  claim 19 , wherein determining the fusion feature by converting the initial feature into the initial feature that has the dimension matching the text encoding comprises:
 determining, based on the initial feature, a key feature and a value feature for an attention mechanism; and   determining, based on the key feature, the value feature, and a predetermined query feature, the fusion feature with the attention mechanism.

Join the waitlist — get patent alerts

Track US2026065525A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.