US2015161179A1PendingUtilityA1
Automatic determination of whether a document includes an image gallery
Est. expiryJun 21, 2024(expired)· nominal 20-yr term from priority
G06F 16/532G06F 18/24G06F 17/30277G06F 17/30864G06F 16/58G06F 16/951G06F 16/9538
56
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Image galleries are automatically located within documents, such as web pages. Documents that are determined to contain image galleries may be treated differently when storing the document for later retrieval by an image search engine. In one implementation, the image galleries are automatically located within a document by calculating position information indicating relative positions of images in the document. The document may be determined to contain an image gallery when the position information indicates that the images in the document are generally evenly distributed.
Claims
exact text as granted — not AI-modified1 - 35 . (canceled)
36 . A non-transitory computer-readable medium storing instructions, the instructions comprising:
one or more instructions which, when executed by one or more processors, cause the one or more processors to parse a document for information identifying a plurality of elements included in the document,
the plurality of elements including one or more words and visual content;
one or more instructions which, when executed by the one or more processors, cause the one or more processors to populate, based on the information identifying the plurality of elements, a data structure that includes information identifying distances between the plurality of elements in the document; one or more instructions which, when executed by the one or more processors, cause the one or more processors to determine, based on the data structure, a distance between:
a word of the one or more words, and
particular visual content of the visual content; and
one or more instructions which, when executed by the one or more processors, cause the one or more processors to determine how the word is related to the particular visual content based on the distance between the word and the particular visual content.
37 . The non-transitory computer-readable medium of claim 36 , where the document includes a web document, and
where the information identifying the plurality of elements includes tags identifying the plurality of elements.
38 . The non-transitory computer-readable medium of claim 37 , where the one or more instructions to parse the document include:
one or more instructions which, when executed by the one or more processors, cause the one or more processors to parse the web document for the tags identifying the plurality of elements.
39 . The non-transitory computer-readable medium of claim 36 , where the one or more instructions to populate the data structure include:
one or more instructions which, when executed by the one or more processors, cause the one or more processors to assign coordinate values, of cells of the data structure, to the information identifying the plurality of elements.
40 . The non-transitory computer-readable medium of claim 39 , where the one or more instructions to determine the distance between the word and the particular visual content include:
one or more instructions which, when executed by the one or more processors, cause the one or more processors to determine the distance between the word and the particular visual content based on particular coordinate values, of the coordinate values, associated with the word and the particular visual content.
41 . The non-transitory computer-readable medium of claim 39 , the instructions further comprising:
one or more instructions which, when executed by the one or more processors, cause the one or more processors to estimate a layout of the document based on the coordinate values; and one or more instructions which, when executed by the one or more processors, cause the one or more processors to analyze content of the document based on the layout of the document.
42 . The non-transitory computer-readable medium of claim 41 , where the one or more instructions to estimate the layout of the document include:
one or more instructions which, when executed by the one or more processors, cause the one or more processors to estimate a geometric layout of the document based on the coordinate values.
43 . A method comprising:
parsing, by one or more processors, a document for information identifying a plurality of elements included in the document,
the plurality of elements including one or more words and visual content;
populating, by the one or more processors and based on the information identifying the plurality of elements, a data structure that includes information identifying distances between the plurality of elements in the document; determining, by the one or more processors and based on the data structure, a distance between:
a word of the one or more words, and
particular visual content of the visual content; and
determining, by the one or more processors, whether the word is related to the particular visual content based on the distance between the word and the particular visual content.
44 . The method of claim 43 , where the document includes a web document, and
where the information identifying the plurality of elements includes tags identifying the plurality of elements.
45 . The method of claim 44 , where parsing the document includes:
parsing the web document for the tags identifying the plurality of elements.
46 . The method of claim 43 , where populating the data structure includes:
assigning coordinate values, of cells of the data structure, to the information identifying the plurality of elements.
47 . The method of claim 46 , where determining the distance between the word and the particular visual content includes:
determining the distance between the word and the particular visual content based on particular coordinate values, of the coordinate values, associated with the word and the particular visual content.
48 . The method of claim 46 , further comprising:
estimating a geometric layout of the document based on the coordinate values; and analyzing content of the document based on the geometric layout of the document.
49 . The method of claim 43 , where the particular visual content corresponds to an image, and
where determining whether the word is related to the particular visual content includes:
determining whether the word is related to the image based on the distance between the word and the image.
50 . A system comprising:
one or more processors to:
parse a document for information identifying a plurality of elements included in the document,
the plurality of elements including one or more words and visual content;
populate, based on the information identifying the plurality of elements, a data structure that includes information identifying distances between the plurality of elements in the document;
determine, based on the data structure, a distance between:
a word of the one or more words, and
particular visual content of the visual content; and
determine that the word is related to the particular visual content based on the distance between the word and the particular visual content.
51 . The system of claim 50 , where the document includes a web document.
52 . The system of claim 51 , where the information identifying the plurality of elements includes tags identifying the plurality of elements, and
where, when parsing the document, the one or more processors are to parse the web document for the tags identifying the plurality of elements.
53 . The system of claim 50 , where, when populating the data structure, the one or more processors are to:
assign coordinate values, of cells of the data structure, to the information identifying the plurality of elements.
54 . The system of claim 53 , where, when determining the distance between the word and the particular visual content, the one or more processors are to:
determine the distance between the word and the particular visual content based on the coordinate values.
55 . The system of claim 53 , where the one or more processors are further to:
estimate a geometric layout of the document based on the coordinate values; and analyze content of the document based on the geometric layout of the document.Join the waitlist — get patent alerts
Track US2015161179A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.