Automated classification and interpretation of life science documents
Abstract
A computer-implemented tool for automated classification and interpretation of documents, such as life science documents supporting clinical trials, is configured to perform a combination of raw text, document construct, and image analyses to enhance classification accuracy by enabling a more comprehensive machine-based understanding of document content. The combination of analyses provides context for classification by leveraging relative spatial relationships among text and image elements, identifying characteristics and formatting of elements, and extracting additional metadata from the documents as compared to conventional automated classification tools, wherein natural language processing (NLP) is applied to associate text with tokens, and relevant differences and similarities between protocols are identified.
Claims
exact text as granted — not AI-modifiedWhat is claimed:
1 . A computer-implemented method, the method comprising:
receiving a plurality of life science documents from one or more databases; applying a machine-learning system to text and images within each of the life science documents, wherein the machine-learning system identifies clinical trials within the life science documents, wherein metadata is identified from the clinical trials, and wherein protocol content based on the identified metadata from the clinical trials is identified; grouping the protocol content into clusters based on one or more similarities; and transferring the clusters of the grouped protocol content into a natural language processing (NLP) database.
2 . The computer-implemented method of claim 1 , further comprising:
identifying whether the protocol content is part of a new protocol or a protocol that has been previously used.
3 . The computer-implemented method of claim 1 , further comprising:
identifying one or more risk factors associated with the clinical trials.
4 . The computer-implemented method of claim 1 , further comprising:
identifying patient burdens associated with patients within the clinical trials.
5 . The computer-implemented method of claim 1 , further comprising:
identifying differences between protocols based on the clusters of the grouped protocol content.
6 . The computer-implemented method of claim 1 , further comprising:
identifying one or more algorithms associated with the clinical trials.
7 . The computer-implemented method of claim 1 , further comprising:
outputting clinical trial data onto a display of the computer.
8 . A computer program product comprising a tangible storage medium encoded with processor-readable instructions that, when executed by one or more processors, enable the computer program product to:
receive a plurality of life science documents from one or more databases; apply a machine-learning system to text and images within each of the life science documents, wherein the machine-learning system identifies clinical trials within the life science documents, wherein metadata is identified from the clinical trials, and wherein protocol content based on the identified metadata from the clinical trials is identified; group the protocol content into clusters based on one or more similarities; and transfer the clusters of the grouped protocol content into a natural language processing (NLP) database.
9 . The computer program product of claim 8 , wherein the machine-learning system identifies whether the protocol content is part of a protocol that has been previously used.
10 . The computer program product of claim 8 , wherein unexpected data is identified by applying a risk assessment onto the clinical trials.
11 . The computer program product of claim 8 , wherein one or more risk factors for patients are identified.
12 . The computer program product of claim 8 , wherein one or more levels of the risks from the clinical trials are identified.
13 . The computer program product of claim 8 , wherein one or more biomarkers associated with increased levels of risk are identified.
14 . The computer program product of claim 8 , wherein regulatory factors associated with the clinical trials are identified.
15 . A computer system connected to a network, the system comprising:
a memory configured to store instructions; one or more processors configured to execute the instructions to perform operations to: receive a plurality of life science documents from one or more databases; apply a machine-learning system to text and images within each of the life science documents, wherein the machine-learning system identifies clinical trials within the life science documents, wherein metadata is identified from the clinical trials, and wherein protocol content based on the identified metadata from the clinical trials is identified; group the protocol content into clusters based on one or more similarities; and transfer the clusters of the grouped protocol content into a natural language processing (NLP) database.
16 . The system of claim 15 , the machine-learning system identifies if the protocol content is part of a new protocol within the clinical trials.
17 . The system of claim 15 , wherein the machine-learning system identifies one or more algorithms within the protocol content.
18 . The system of claim 15 , wherein clinical trial data associated with one or more treatment is identified.
19 . The system of claim 15 , wherein areas of risk within the protocol content is identified.
20 . The system of claim 15 , wherein changes to the protocol content are identified.Join the waitlist — get patent alerts
Track US2023177267A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.