US2015234917A1PendingUtilityA1
Per-document index for semantic searching
Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Nov 28, 2012Filed: Apr 30, 2015Published: Aug 20, 2015
Est. expiryNov 28, 2032(~6.3 yrs left)· nominal 20-yr term from priority
G06F 17/30622G06F 17/30864G06F 16/319G06F 16/951G06F 16/31
47
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Methods, computer systems, and computer-storage medium for generating a per-document index used for semantic searching is provided. A document is received and parsed into a plurality of section. Each term in each section is translated in order to at least one of a cache index or a term identifier. Subsequent to translating the terms, each section is separately group encoded to generate the per-document index. The per-document index is stored in association with a data store.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . One or more computer-storage media storing computer-useable instructions that, when used by one or more computing devices, cause the one or more computing devices to perform a method of generating a per-document index for semantic searching, the method comprising:
receiving a document, the document comprising the original document and annotations associated with the document; parsing the document into a plurality of sections; for each section of the plurality of sections, translating in order each term to at least one of its corresponding cache index, or its corresponding term identifier; subsequent to translating in order the each term, group encoding the each section of the plurality of sections, the group-encoded sections comprising a per-document index for the document; and storing the per-document index in a data store.
2 . The media of claim 1 , further comprising:
generating a custom-dictionary section for the document, the custom-dictionary section comprising document-specific terms at a specified position within the custom-dictionary section; group encoding the custom-dictionary section; and storing the encoded custom-dictionary section is association with the per-document index.
3 . The media of claim 2 , wherein parsing the document further comprises identifying one or more attributes associated with the document and generating meta-word terms for each of the one or more attributes.
4 . The media of claim 3 , wherein the one or more attributes comprise at least one or more of location information associated with the document, originating language of the document, and fingerprint information associated with the document.
5 . The media of claim 3 , wherein each of the meta-word terms includes a prefix identifying a type of attribute.
6 . The media of claim 5 , wherein the each of the meta-word terms further includes a value associated with prefix.
7 . The media of claim 2 , wherein the plurality of sections comprise at least a document data section, a body section, one or more meta-word sections, and one or more meta-stream sections.
8 . A computerized method carried out by at least one server having at least one processor for generating a per-document index for semantic searching, the method comprising:
receiving a document, the document comprising the original document and annotations associated with the document; parsing the document into a plurality of sections; for each section of the plurality of sections, translating, using the at least one processor, in order each term to at least one of its corresponding cache index, or its corresponding term identifier; subsequent to translating in order the each term, group encoding the each section of the plurality of sections, the group-encoded sections comprising a per-document index for the document; and storing the per-document index in a data store.
9 . The method of claim 8 , further comprising:
generating a custom-dictionary section for the document, the custom-dictionary section comprising document-specific terms at a specified position within the custom-dictionary section; group encoding the custom-dictionary section; and storing the encoded custom-dictionary section is association with the per-document index.
10 . The method of claim 9 , wherein translating in order the each term to its corresponding term identifier comprises:
accessing at least one of a section-specific dictionary or the custom-dictionary section; and identifying a position of the each term in the at least one of the section-specific dictionary or the custom-dictionary section, the position of the each term comprising the each term's term identifier.
11 . The method of claim 10 , wherein the plurality of sections comprise at least a document data section, a body section, one or more meta-word sections, and one or more meta-stream sections.
12 . The method of claim 11 , wherein the body section and the one or more meta-streams sections share the same section-specific dictionary.
13 . The method of claim 8 , wherein translating in order the each term to its corresponding cache index comprises processing the each term's corresponding term identifier through an entry cache.
14 . The method of claim 8 , wherein the term identifier is a numerical value.
15 . The method of claim 14 , wherein the numerical value is smaller when the each term is commonly used, and wherein the numerical value is larger when the each term is infrequently used.
16 . A system for generating a per-document index for semantic searching, the system comprising:
a server having one or more processors and one or more computer-storage media; a data store coupled with the server, wherein the server:
receives a document, the document comprising the original document and annotations associated with the document;
parses the document into a plurality of sections;
for each section of the plurality of sections, translates in order each term to at least one of its corresponding cache index, or its corresponding term identifier;
subsequent to translating in order the each term, group encodes the each section of the plurality of sections, the group-encoded sections comprising a per-document index for the document; and
stores the per-document index in the data store.
17 . The system of claim 16 , wherein the plurality of sections comprise at least a document data section, a body section, one or more meta-word sections, and one or more meta-stream sections.
18 . The system of claim 16 , wherein the server further:
generates a custom-dictionary section for the document, the custom-dictionary section comprising document-specific terms at a specified position within the custom-dictionary section; group encodes the custom-dictionary section; and stores the encoded custom-dictionary section is association with the per-document index in the data store.
19 . The system of claim 16 , wherein the each term of the each section is associated with one or more of a position of the each term within the section, the term identifier, and metadata associated with the each term.
20 . The system of claim 16 , wherein the data store comprises a solid state drive (SSD).Join the waitlist — get patent alerts
Track US2015234917A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.