Indexing and search query processing
Abstract
A method for processing a search query according to one embodiment includes receiving a search query containing terms; combining at least some consecutive terms in the search query to create biwords; looking up at least some of the terms and biwords in a search index for identifying sections of documents containing the at least some of the terms and/or biwords; generating a content score for each of the identified sections based at least in part on a number of the terms and biwords found in the sections of each document, wherein the biwords are given a higher priority than matched terms, wherein the priority affects the content score; and selecting and outputting an indicator of at least one of the sections, or portion thereof, based at least in part on the content score.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A non-transitory computer readable storage medium having an inverted index structure stored thereon for access by a program configured to execute keyword searches, the inverted index structure comprising:
a plurality of content words extracted from an unstructured text document; for each of the content words, at least one document identifier containing information about the unstructured text document containing the content word; and for each of the document identifiers, at least one position identifier containing first information about a section in the unstructured text document containing the content word and second information about a paragraph in the section in the unstructured text document containing the content word, the first information being distinct from the second information.
2 . The inverted index structure as recited in claim 1 , wherein at least some of the position identifiers further contain information about a page in the unstructured text document containing the content word.
3 . The inverted index structure as recited in claim 1 , wherein at least some of the position identifiers include a weighting value of the content word.
4 . The inverted index structure as recited in claim 3 , wherein the weighting value is based at least in part on a position of the content word in the unstructured text document.
5 . The inverted index structure as recited in claim 1 , further comprising context meta data associated with the unstructured text document, the context meta data indicating a context of the unstructured text document associated therewith.
6 . The inverted index structure as recited in claim 5 , wherein at least some of the context meta data is weighted.
7 . The inverted index structure as recited in claim 1 , wherein for at least one of the position identifiers, the first information is stored within a multi-bit integer of which a first plurality of bits store a section identifier associated with the respective section and a second plurality of bits store at least one priority bit.Join the waitlist — get patent alerts
Track US2019026300A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.