Information object extraction using combination of classifiers
Abstract
Systems and methods for information extraction from natural language texts using a combination of classifier models. An example method may comprise: producing, by performing syntactico-semantic analysis of a natural language text, a plurality of syntactico-semantic structures representing the natural language text; identifying, using a first classifier model to process a first plurality of classification attributes derived from the syntactico-semantic structures, a plurality of core constituents, such that each core constituent of the plurality of core constituents is associated with a span of a plurality of spans, wherein each span represents an attribute of an information object of a specified ontology class; identifying, using a second classifier model to process a second plurality of classification attributes derived from the syntactico-semantic structures, child constituents of each of the plurality of core constituents; and determining, using a third classifier model to process a third plurality of classification attributes derived from the syntactico-semantic structures, whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
processing, by a computer system, based on a first classifier model, a first plurality of classification attributes derived from a natural language text, wherein the first classifier model identifies a plurality of core constituents, wherein each core constituent is associated with an information object of a specified ontology class; identifying a plurality of spans, wherein each span of the plurality of spans includes a core constituent of the plurality of core constituents and further includes one or more child constituents of the core constituent; and processing, based on a second classifier model, a second plurality of classification attributes derived from the natural language text, wherein the second classifier model determines whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.
2 . The method of claim 1 , further comprising:
utilizing, for performing a natural language processing task, the information object attributes associated with the first span and the second span.
3 . The method of claim 1 , further comprising:
displaying, in visual association with a first projection of the first span in the natural language text and a second projection of the second span in the natural language text, the information object attributes associated with the first span and the second span; and accepting user input to perform at least one of: confirming the information object attributes or modifying the information object attributes.
4 . The method of claim 1 , further comprising:
determining, using a training data set, a parameter of the first classifier model, wherein the training data set comprises an annotated natural language text comprising a plurality of textual annotations, wherein each textual annotation is associated with an information object attribute of an information object of a known category.
5 . The method of claim 1 , wherein the first classifier model yields a likelihood of a candidate node being a core constituent of a span that represents an attribute of an information object of the specified ontology class.
6 . The method of claim 1 , wherein the first plurality of classification attributes include attributes of a candidate core constituent and at least one of: a parent node of the candidate core constituent, a child node of the candidate core constituent, or a sibling node of the candidate core constituent.
7 . The method of claim 1 , wherein the second classifier model yields a likelihood of the first span and the second span being associated with the same information object.
8 . The method of claim 1 , wherein the second plurality of classification attributes include attributes of nodes of the first span and attributes of nodes of the second span.
9 . A system, comprising:
a memory; a processor, coupled to the memory, the processor configured to:
process, based on a first classifier model, a first plurality of classification attributes derived from a natural language text, wherein the first classifier model identifies a plurality of core constituents, wherein each core constituent is associated with an information object of a specified ontology class;
identify a plurality of spans, wherein each span of the plurality of spans includes a core constituent of the plurality of core constituents and further includes one or more child constituents of the core constituent; and
process, based on a second classifier model, a second plurality of classification attributes derived from the natural language text, wherein the second classifier model determines whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.
10 . The system of claim 9 , wherein the processor is further configured to:
utilize, for performing a natural language processing task, the information object attributes associated with the first span and the second span.
11 . The system of claim 9 , wherein the processor is further configured to:
display, in visual association with a first projection of the first span in the natural language text and a second projection of the second span in the natural language text, the information object attributes associated with the first span and the second span; and accept user input to perform at least one of: confirming the information object attributes or modifying the information object attributes.
12 . The system of claim 9 , wherein the processor is further configured to:
determine, using a training data set, a parameter of the first classifier model, wherein the training data set comprises an annotated natural language text comprising a plurality of textual annotations, wherein each textual annotation is associated with an information object attribute of an information object of a known category.
13 . The system of claim 9 , wherein the first classifier model yields a likelihood of a candidate node being a core constituent of a span that represents an attribute of an information object of the specified ontology class.
14 . The system of claim 9 , wherein the first plurality of classification attributes include attributes of a candidate core constituent and at least one of: a parent node of the candidate core constituent, a child node of the candidate core constituent, or a sibling node of the candidate core constituent.
15 . The system of claim 9 , wherein the second classifier model yields a likelihood of the first span and the second span being associated with the same information object.
16 . The system of claim 9 , wherein the second plurality of classification attributes include attributes of nodes of the first span and attributes of nodes of the second span.
17 . A computer-readable non-transitory storage medium comprising executable instructions that, when executed by a computer system, cause the computer system to:
process, based on a first classifier model, a first plurality of classification attributes derived from a natural language text, wherein the first classifier model identifies a plurality of core constituents, wherein each core constituent is associated with an information object of a specified ontology class; identify a plurality of spans, wherein each span of the plurality of spans includes a core constituent of the plurality of core constituents and further includes one or more child constituents of the core constituent; and process, based on a second classifier model, a second plurality of classification attributes derived from the natural language text, wherein the second classifier model determines whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.
18 . The computer-readable non-transitory storage medium of claim 17 , further comprising executable instructions that, when executed by the computer system, cause the computer system to:
utilize, for performing a natural language processing task, the information object attributes associated with the first span and the second span.
19 . The computer-readable non-transitory storage medium of claim 17 , further comprising executable instructions that, when executed by the computer system, cause the computer system to:
display, in visual association with a first projection of the first span in the natural language text and a second projection of the second span in the natural language text, the information object attributes associated with the first span and the second span; and accept user input to perform at least one of: confirming the information object attributes or modifying the information object attributes.
20 . The computer-readable non-transitory storage medium of claim 17 , further comprising executable instructions that, when executed by the computer system, cause the computer system to:
determine, using a training data set, a parameter of the first classifier model, wherein the training data set comprises an annotated natural language text comprising a plurality of textual annotations, wherein each textual annotation is associated with an information object attribute of an information object of a known category.Join the waitlist — get patent alerts
Track US2020057810A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.