Extracting facts from unstructured data
Abstract
Methods, systems, and apparatus, including computer programs encoded on computer storage media, to present a video. One of the methods includes obtaining one or more unstructured documents. The method includes obtaining, by a computer system, a data model, the data model identifying a type of fact that can be determined from the one or more unstructured documents. The method includes determining, by the computer system, a channel to extract facts from the document based on the type of fact. The method includes distributing, by the computer system, the one or more unstructured documents to the channel. The method includes extracting, by the channel, facts from the one or more unstructured documents. The method also includes storing the facts in a data model.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method for processing data originating from electronic medical records, comprising:
obtaining one or more unstructured documents; obtaining, by a computer system, a data model, the data model identifying a type of fact that can be determined from the one or more unstructured documents; determining, by the computer system, a channel to extract facts from the document based on the type of fact; distributing, by the computer system, the one or more unstructured documents to the channel; extracting, by the channel, facts from the one or more unstructured documents; and storing the facts in a data model.
2 . The computer-implemented method of claim 1 , wherein extracting facts is performed by a computer system.
3 . The computer-implemented method of claim 1 , further comprising verifying the facts stored in the data model, wherein the one or more unstructured documents include training documents for which the facts are known.
4 . The computer-implemented method of claim 1 , further comprising identifying a cohort of patients based on the extracted facts.
5 . The computer-implemented method of claim 1 , further comprising:
determining a measure of quality for the extracted facts; and updating the data model based on the measure of quality.
6 . The computer-implemented method of claim 1 , further comprising:
identifying an extracted fact as being longitudinal; comparing the extracted fact to previously extracted facts of the same type for the same patient.
7 . The computer-implemented method of claim 1 , further comprising:
extracting a set of facts from a set of unstructured documents; establishing the set of unstructured documents as a training set; training a model using the set of facts and the set of unstructured documents; and extracting new facts from new unstructured documents using the model.
8 . A system comprising:
one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
obtaining one or more unstructured documents;
obtaining, by a computer system, a data model, the data model identifying a type of fact that can be determined from the one or more unstructured documents;
determining, by the computer system, a channel to extract facts from the document based on the type of fact;
distributing, by the computer system, the one or more unstructured documents to the channel;
extracting, by the channel, facts from the one or more unstructured documents; and
storing the facts in a data model.
9 . (canceled)
10 . The system of claim 8 , wherein the operations further comprise verifying the facts stored in the data model, wherein the one or more unstructured documents include training documents for which the facts are known.
11 . The system of claim 8 , wherein the operations further comprise identifying a cohort of patients based on the extracted facts.
12 . The system of claim 8 , wherein the operations further comprise:
determining a measure of quality for the extracted facts; and updating the data model based on the measure of quality.
13 . The system of claim 8 , wherein the operations further comprise:
identifying an extracted fact as being longitudinal; comparing the extracted fact to previously extracted facts of the same type for the same patient.
14 . The system of claim 8 , wherein the operations further comprise:
extracting a set of facts from a set of unstructured documents; establishing the set of unstructured documents as a training set; training a model using the set of facts and the set of unstructured documents; and extracting new facts from new unstructured documents using the model.
15 . A non-transitory computer storage medium encoded with computer program instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:
obtaining one or more unstructured documents; obtaining, by a computer system, a data model, the data model identifying a type of fact that can be determined from the one or more unstructured documents; determining, by the computer system, a channel to extract facts from the document based on the type of fact; distributing, by the computer system, the one or more unstructured documents to the channel; extracting, by the channel, facts from the one or more unstructured documents; and storing the facts in a data model.
16 . The non-transitory computer storage medium of claim 15 , wherein extracting facts is performed by a computer system.
17 . The non-transitory computer storage medium of claim 15 , wherein the operations further comprise verifying the facts stored in the data model, wherein the one or more unstructured documents include training documents for which the facts are known.
18 . The non-transitory computer storage medium of claim 15 , wherein the operations further comprise identifying a cohort of patients based on the extracted facts.
19 . The non-transitory computer storage medium of claim 15 , wherein the operations further comprise:
determining a measure of quality for the extracted facts; and updating the data model based on the measure of quality.
20 . The non-transitory computer storage medium of claim 15 , wherein the operations further comprise:
identifying an extracted fact as being longitudinal; comparing the extracted fact to previously extracted facts of the same type for the same patient.
21 . The non-transitory computer storage medium of claim 15 , wherein the operations further comprise:
extracting a set of facts from a set of unstructured documents; establishing the set of unstructured documents as a training set; training a model using the set of facts and the set of unstructured documents; and extracting new facts from new unstructured documents using the model.Join the waitlist — get patent alerts
Track US2020410400A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.