US2026004052A1PendingUtilityA1

Image context based text generation

Assignee: SHUTTERSTOCK INCPriority: Jun 26, 2024Filed: Jun 26, 2024Published: Jan 1, 2026
Est. expiryJun 26, 2044(~17.9 yrs left)· nominal 20-yr term from priority
G06V 20/70G06F 40/106G06F 16/532G06F 40/166G06F 16/33295
41
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, systems, and storage media for generating contextually relevant text from image descriptions and user intent are disclosed. Exemplary implementations may: receive an image and a user-defined intent for text output; analyze the received image to generate a contextual description of the image; generate a query based on the contextual description of the image and the user-defined intent; and generate the text output based on the query.

Claims

exact text as granted — not AI-modified
1 . A method for generating contextually relevant text from image descriptions and user intent, comprising:
 receiving an image and a user-defined intent for text output, the user-defined intent including at least a description of an intended use, a selected tone, and a selected content type for the text output;   generating a system message including at least a first text argument based on the user-defined intent for the text output;   analyzing the received image to generate a contextual description of the image;   modifying, in response to a selection for an image influence feature, the system message by prepending a second text argument including the contextual description of the image to the first text argument, wherein the image influence feature enables consideration of the image;   generating a query based on the system message; and   generating the text output based on the query, wherein the text output is formatted and structured according to a set of guidelines corresponding to the selected content type, and contents of the text output align with a semantic, stylistic, and thematic characteristics of the image.   
     
     
         2 . The method of  claim 1 , further comprising generating a notification on a user device inviting the user to input the user-defined intent for the text output immediately after the user downloads the image. 
     
     
         3 . The method of  claim 1 , wherein receiving the selection for the image influence feature further comprises providing a toggle option within a text generator interface that allows the user to activate or deactivate the image influence feature, wherein activating the image influence feature enables consideration of the contextual description of the image in the text output generation and deactivating the image influence feature disables consideration of the contextual description of the image in the text output generation. 
     
     
         4 . The method of  claim 1 , further comprising embedding the generated text output with metadata that includes the contextual description of the image and the user-defined intent to facilitate indexing and retrieval of the text output. 
     
     
         5 . The method of  claim 1 , further comprising presenting a preview of the text output in a layout corresponding to the selected content type on a text generator interface for user confirmation before finalizing the text output. 
     
     
         6 . The method of  claim 1 , further comprising filtering the user-defined intent through a safety and privacy guideline checker to detect and respond to potentially offensive terms before generating the text output. 
     
     
         7 . The method of  claim 1 , wherein the analyzing of the received image includes:
 extracting visual elements such as colors, objects, and activities to enhance the contextual description; and   applying at least a portion of the visual elements to the text output.   
     
     
         8 . The method of  claim 1 , wherein the generated text output is configured for dissemination across multiple social media platforms, each platform receiving a customized version of the text output. 
     
     
         9 . The method of  claim 1 , wherein the user-defined intent includes a content type selection, the content type selection including options for various event invitations, and the generated text output is tailored to match a specific event theme and audience based on the content type selection. 
     
     
         10 . The method of  claim 1 , wherein the selected tone is selected from tone options within a text generator interface, and the selected tone is dynamically adjusted in response to real-time sentiment analysis of the user-defined intent and the contextual description of the image. 
     
     
         11 . A system configured for generating contextually relevant text from image descriptions and user intent, the system comprising:
 one or more hardware processors configured by machine-readable instructions to: receive an image and a user-defined intent for text output, the user-defined intent including at least a description of an intended use, a selected tone, and a selected content type for the text output;   generate a system message including at least a first text argument based on the user-defined intent for the text output;   analyze the received image to generate a contextual description of the image;   modify, in response to a selection for an image influence feature, the system message by prepending a second text argument including the contextual description of the image to the first text argument, wherein the image influence feature enables consideration of the image;   generate a query based on the system message; and   generate the text output based on the query, wherein the text output is formatted and structured according to a set of guidelines corresponding to the selected content type, and contents of the text output align with a semantic, stylistic, and thematic characteristics of the image.   
     
     
         12 . The system of  claim 11 , wherein the one or more hardware processors are further configured by machine-readable instructions to generate a notification on a user device inviting the user to input the user-defined intent for the text output immediately after the user downloads the image. 
     
     
         13 . The system of  claim 11 , wherein the one or more hardware processors are further configured by machine-readable instructions to provide a toggle option within a text generator interface that allows the user to activate or deactivate the image influence feature, wherein activating the image influence feature enables consideration of the contextual description of the image in the text output generation and deactivating the image influence feature disables consideration of the contextual description of the image in the text output generation. 
     
     
         14 . The system of  claim 11 , wherein the one or more hardware processors are further configured by machine-readable instructions to embed the generated text output with metadata that includes the contextual description of the image and the user-defined intent to facilitate indexing and retrieval of the text output. 
     
     
         15 . The system of  claim 11 , wherein the one or more hardware processors are further configured by machine-readable instructions to present a preview of the text output in a layout corresponding to the selected content type on a text generator interface for user confirmation before finalizing the text output. 
     
     
         16 . The system of  claim 11 , wherein the one or more hardware processors are further configured by machine-readable instructions to filter the user-defined intent through a safety and privacy guideline checker to detect and respond to potentially offensive terms before generating the text output. 
     
     
         17 . The system of  claim 11 , wherein the analyzing of the received image includes:
 extracting visual elements such as colors, objects, and activities to enhance the contextual description; and   applying at least a portion of the visual elements to the text output.   
     
     
         18 . The system of  claim 11 , wherein the generated text output is configured for dissemination across multiple social media platforms, each platform receiving a customized version of the text output. 
     
     
         19 . The system of  claim 11 , wherein the user-defined intent includes a content type selection, the content type selection including options for various event invitations, and the generated text output is tailored to match a specific event theme and audience based on the content type selection. 
     
     
         20 . A non-transient computer-readable storage medium having instructions embodied thereon, the instructions being executable by one or more processors to perform a method for generating contextually relevant text from image descriptions and user intent, the method comprising:
 receiving an image and a user-defined intent for text output, the user-defined intent including at least a description of an intended use, a selected tone, and a selected content type for the text output;   generating a system message including at least a first text argument based on the user-defined intent for the text output;   analyzing the received image to generate a contextual description of the image;   modifying, in response to a selection for an image influence feature, the system message by prepending a second text argument including the contextual description to the first text argument, wherein the image influence feature enables consideration of the image;   generating a query based on the system message; and   generating the text output based on the query, wherein the text output is formatted and structured according to a set of guidelines corresponding to the selected content type, and contents of the text output align with a semantic, stylistic, and thematic characteristics of the image.

Join the waitlist — get patent alerts

Track US2026004052A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.