US2022075806A1PendingUtilityA1

Natural language image search

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: May 16, 2014Filed: Nov 17, 2021Published: Mar 10, 2022
Est. expiryMay 16, 2034(~7.8 yrs left)· nominal 20-yr term from priority
G06F 16/53G06F 16/55G06F 40/00G06F 16/50G06F 16/5866G06F 16/285G06F 16/3329G06F 16/9024
55
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Natural language image search is described, for example, whereby natural language queries may be used to retrieve images from a store of images automatically tagged with image tags being concepts of an ontology (which may comprise a hierarchy of concepts). In various examples, a natural language query is mapped to one or more of a plurality of image tags, and the mapped query is used for retrieval. In various examples, the query is mapped by computing one or more distance measures between the query and the image tags, the distance measures being computed with respect to the ontology and/or with respect to a semantic space of words computed from a natural language corpus. In examples, the image tags may be associated with bounding boxes of objects depicted in the images, and a user may navigate the store of images by selecting a bounding box and/or an image.

Claims

exact text as granted — not AI-modified
1 . A computer-implemented method comprising:
 receiving a natural language query;   computing a first distance in an ontology between the natural language query and individual ones of a plurality of image tags, each image tag being a concept of the ontology;   computing at least one second distance in a semantic space of words between the natural language query and individual ones of the plurality of image tags;   selecting at least one of the plurality of image tags on the basis of the computed first and second distances; and   retrieving, using the selected at least one image tag, one or more images from a database of images tagged with the selected image tags.   
     
     
         2 . The method of  claim 1 , wherein the first distance is computed by traversing between nodes in the ontology, wherein the ontology is a graph of nodes representing concepts, the nodes being linked by edges according to relationships between the concepts. 
     
     
         3 . The method of  claim 1 , wherein the semantic space of words has been learnt from a corpus of natural language documents. 
     
     
         4 . The method of  claim 3 , wherein the semantic space of words has been learnt using a neural network. 
     
     
         5 . The method of  claim 1 , wherein one of the at least one second distance is computed using a distance metric selected from any of: cosine similarity, dot product, dice similarity, hamming distance, and city block distance. 
     
     
         6 . The method of  claim 1 , wherein computing at least one second distance comprises computing at least two second distances. 
     
     
         7 . The method of  claim 1 , wherein selecting at least one of the plurality of image tags on the basis of the computed first and second distances comprises ignoring any of the first and second distances that exceed a predetermined threshold. 
     
     
         8 . The method of  claim 1 , wherein selecting at least one of the plurality of image tags on the basis of the computed first and second distances comprises:
 representing each computed first and second distance as a vote for a particular image tag of the plurality of image tags;   combining the votes for each image tag;   selecting one or more image tags based on the number of votes.   
     
     
         9 . The method of  claim 8 , wherein each vote is weighted based on the magnitude of the corresponding distance prior to combining the votes. 
     
     
         10 . The method of  claim 1 , further comprising:
 displaying at least a portion of the one or more retrieved images;   receiving information indicating one of the retrieved images has been selected; and   displaying the selected image and information related to the selected image.   
     
     
         11 . The method of  claim 10 , wherein the information related to the selected image comprises one or more images that are similar to the selected image. 
     
     
         12 . The method of  claim 11 , wherein the similarity of two images is based on image tags shared between the two images and confidence values associated with each shared tag. 
     
     
         13 . The method of  claim 10 , further comprising:
 receiving information indicating the position of a cursor with respect to the selected image, the cursor being controlled by a user;   determining whether the cursor is positioned over an object identified in the selected image; and   in response to determining the cursor is positioned over an object identified in the selected image, displaying a bounding box around the identified object.   
     
     
         14 . The method of  claim 13 , further comprising:
 receiving an indication that the bounding box has been selected; and   updating the natural language query to include an image tag associated with the identified object corresponding to the bounding box.   
     
     
         15 . The method of  claim 1 , wherein the natural language query comprises a plurality of query terms and an indication of whether the terms are to be proximate, and in response to determining the terms are to be proximate, retrieving one or more images from the database of images tagged with each of the selected image tags wherein objects associated with the selected image tags are proximate. 
     
     
         16 . The method of  claim 1 , further comprising automatically generating the database of tagged images from a plurality of untagged images using one or more trained machine learning components, each trained machine learning component trained to identify one or more features in an image and assign one or more tags to individual identified features. 
     
     
         17 . The method of  claim 1 , further comprising:
 receiving data indicating the one or more retrieved images are to be shared; and   making the one or more retrieved images available to one or more other parties.   
     
     
         18 . A system comprising a computing-based device configured to:
 receive a natural language query;   compute a first distance in an ontology between the natural language query and individual ones of a plurality of image tags, an image tag being a concept of the ontology;   compute at least one second distance in a semantic space of words between the natural language query and individual ones of the plurality of image tags;   select at least one of the plurality of image tags on the basis of the computed first and second distances; and   retrieve, using the selected at least one image tag, one or more images from a database of images tagged with the selected image tags.   
     
     
         19 . The system according to  claim 18 , the computing-based device being at least partially implemented using hardware logic selected from any one of more of: a field-programmable gate array, a program-specific integrated circuit, a program-specific standard product, a system-on-a-chip, a complex programmable logic device. 
     
     
         20 . A computer-implemented method comprising:
 receiving a natural language query;   computing a first distance in an ontology between the natural language query and individual ones of a plurality of image tags, each image tag being a concept of the ontology;   computing at least one second distance in a semantic space of words between the natural language query and individual ones of the plurality of image tags, the semantic space of words being generated by applying a trained neural network to a corpus of natural language documents;   selecting at least one of the plurality of image tags on the basis of the computed first and second distances; and   retrieving, using the selected at least one image tag, one or more images from a database of images tagged with the selected image tags.

Join the waitlist — get patent alerts

Track US2022075806A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.