US2019197061A1PendingUtilityA1

Corpus-scoped annotation and analysis

Assignee: IBMPriority: May 3, 2017Filed: Mar 8, 2019Published: Jun 27, 2019
Est. expiryMay 3, 2037(~10.8 yrs left)· nominal 20-yr term from priority
G06F 16/9024G06F 40/169G06F 16/93G06F 16/2453G06F 17/241G06F 17/27G06F 40/237
63
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Corpus-scoped annotation and analysis. Enrichment analysis data is generated including annotations and metadata for a plurality of documents that are part of a corpus. Whether to generate a second set of annotations is determined, based on a correlation of the annotations and metadata. A relational database is populated with the enrichment analysis data. A corpus-scoped query is resolved, initiated by an application, using the enrichment analysis data and content of the corpus.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer system comprising:
 one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored on at least one of the one or more computer-readable tangible storage devices for execution by at least one of the one or more processors via at least one of the one or more memories, the program instructions comprising:   program instructions to produce a collection of prepared data, from structured and unstructured data, utilizing rule-based and/or machine learning techniques, wherein the structured and unstructured data is provided by a content provider;   program instructions to produce a collection of vetted data, based on the collection of prepared data, utilizing a subject matter expert to create a dictionary of terms;   program instructions to create a trained algorithm based on the collection of vetted data;   program instructions to utilize the trained algorithm in combination with at least one feature from the group consisting of: language identification, semantic scoring, lemmatization, clustering and classification, key phrase extraction, synonym expansion, document and entity sentiment calculation, predictive analytics, and trending entity identification to generate, by the one or more processors, enrichment analysis data, wherein the enrichment analysis data includes annotations and metadata for a plurality of documents that are part of a corpus;   program instructions to determine, by the one or more processors, whether to generate a second set of annotations, based on a correlation of the annotations and metadata;   program instructions to determine, by the one or more processors, whether to generate a second set of annotations, utilizing a re-trained algorithm, wherein the re-trained algorithm is created in response to updates to the collection of vetted data;   program instructions to determine whether to update the enrichment analysis data with the second set of annotations;   program instructions to populate, by the one or more processors, a relational database with the enrichment analysis data, wherein the relational database is configured to manage content of the documents and participate in incremental ingestion during population of the relational database with enrichment analysis data, and wherein the relational database is configured to provide versioning capabilities for the enrichment analysis data stored in the relational database; and   program instructions to resolve, by the one or more processors, a corpus-scoped query initiated by an application using the enrichment analysis data and content of the corpus, wherein resolving the corpus-scoped query comprises optimizing, by the one or more processors, the corpus-scoped query through a query optimization feature of a relational database management system for the relational database.

Join the waitlist — get patent alerts

Track US2019197061A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.