US2020057810A1PendingUtilityA1

Information object extraction using combination of classifiers

Assignee: ABBYY PRODUCTION LLCPriority: Dec 11, 2017Filed: Aug 26, 2019Published: Feb 20, 2020
Est. expiryDec 11, 2037(~11.4 yrs left)· nominal 20-yr term from priority
G06F 40/30G06F 40/211G06F 40/169G06F 17/241G06F 17/2785G06F 17/271G06F 40/40G06F 40/20G06F 40/00
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for information extraction from natural language texts using a combination of classifier models. An example method may comprise: producing, by performing syntactico-semantic analysis of a natural language text, a plurality of syntactico-semantic structures representing the natural language text; identifying, using a first classifier model to process a first plurality of classification attributes derived from the syntactico-semantic structures, a plurality of core constituents, such that each core constituent of the plurality of core constituents is associated with a span of a plurality of spans, wherein each span represents an attribute of an information object of a specified ontology class; identifying, using a second classifier model to process a second plurality of classification attributes derived from the syntactico-semantic structures, child constituents of each of the plurality of core constituents; and determining, using a third classifier model to process a third plurality of classification attributes derived from the syntactico-semantic structures, whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, comprising:
 processing, by a computer system, based on a first classifier model, a first plurality of classification attributes derived from a natural language text, wherein the first classifier model identifies a plurality of core constituents, wherein each core constituent is associated with an information object of a specified ontology class;   identifying a plurality of spans, wherein each span of the plurality of spans includes a core constituent of the plurality of core constituents and further includes one or more child constituents of the core constituent; and   processing, based on a second classifier model, a second plurality of classification attributes derived from the natural language text, wherein the second classifier model determines whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.   
     
     
         2 . The method of  claim 1 , further comprising:
 utilizing, for performing a natural language processing task, the information object attributes associated with the first span and the second span.   
     
     
         3 . The method of  claim 1 , further comprising:
 displaying, in visual association with a first projection of the first span in the natural language text and a second projection of the second span in the natural language text, the information object attributes associated with the first span and the second span; and   accepting user input to perform at least one of: confirming the information object attributes or modifying the information object attributes.   
     
     
         4 . The method of  claim 1 , further comprising:
 determining, using a training data set, a parameter of the first classifier model, wherein the training data set comprises an annotated natural language text comprising a plurality of textual annotations, wherein each textual annotation is associated with an information object attribute of an information object of a known category.   
     
     
         5 . The method of  claim 1 , wherein the first classifier model yields a likelihood of a candidate node being a core constituent of a span that represents an attribute of an information object of the specified ontology class. 
     
     
         6 . The method of  claim 1 , wherein the first plurality of classification attributes include attributes of a candidate core constituent and at least one of: a parent node of the candidate core constituent, a child node of the candidate core constituent, or a sibling node of the candidate core constituent. 
     
     
         7 . The method of  claim 1 , wherein the second classifier model yields a likelihood of the first span and the second span being associated with the same information object. 
     
     
         8 . The method of  claim 1 , wherein the second plurality of classification attributes include attributes of nodes of the first span and attributes of nodes of the second span. 
     
     
         9 . A system, comprising:
 a memory;   a processor, coupled to the memory, the processor configured to:
 process, based on a first classifier model, a first plurality of classification attributes derived from a natural language text, wherein the first classifier model identifies a plurality of core constituents, wherein each core constituent is associated with an information object of a specified ontology class; 
 identify a plurality of spans, wherein each span of the plurality of spans includes a core constituent of the plurality of core constituents and further includes one or more child constituents of the core constituent; and 
 process, based on a second classifier model, a second plurality of classification attributes derived from the natural language text, wherein the second classifier model determines whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object. 
   
     
     
         10 . The system of  claim 9 , wherein the processor is further configured to:
 utilize, for performing a natural language processing task, the information object attributes associated with the first span and the second span.   
     
     
         11 . The system of  claim 9 , wherein the processor is further configured to:
 display, in visual association with a first projection of the first span in the natural language text and a second projection of the second span in the natural language text, the information object attributes associated with the first span and the second span; and   accept user input to perform at least one of: confirming the information object attributes or modifying the information object attributes.   
     
     
         12 . The system of  claim 9 , wherein the processor is further configured to:
 determine, using a training data set, a parameter of the first classifier model, wherein the training data set comprises an annotated natural language text comprising a plurality of textual annotations, wherein each textual annotation is associated with an information object attribute of an information object of a known category.   
     
     
         13 . The system of  claim 9 , wherein the first classifier model yields a likelihood of a candidate node being a core constituent of a span that represents an attribute of an information object of the specified ontology class. 
     
     
         14 . The system of  claim 9 , wherein the first plurality of classification attributes include attributes of a candidate core constituent and at least one of: a parent node of the candidate core constituent, a child node of the candidate core constituent, or a sibling node of the candidate core constituent. 
     
     
         15 . The system of  claim 9 , wherein the second classifier model yields a likelihood of the first span and the second span being associated with the same information object. 
     
     
         16 . The system of  claim 9 , wherein the second plurality of classification attributes include attributes of nodes of the first span and attributes of nodes of the second span. 
     
     
         17 . A computer-readable non-transitory storage medium comprising executable instructions that, when executed by a computer system, cause the computer system to:
 process, based on a first classifier model, a first plurality of classification attributes derived from a natural language text, wherein the first classifier model identifies a plurality of core constituents, wherein each core constituent is associated with an information object of a specified ontology class;   identify a plurality of spans, wherein each span of the plurality of spans includes a core constituent of the plurality of core constituents and further includes one or more child constituents of the core constituent; and   process, based on a second classifier model, a second plurality of classification attributes derived from the natural language text, wherein the second classifier model determines whether a first span of the plurality of spans and a second span of the plurality of spans represent information object attributes that are associated with a same information object.   
     
     
         18 . The computer-readable non-transitory storage medium of  claim 17 , further comprising executable instructions that, when executed by the computer system, cause the computer system to:
 utilize, for performing a natural language processing task, the information object attributes associated with the first span and the second span.   
     
     
         19 . The computer-readable non-transitory storage medium of  claim 17 , further comprising executable instructions that, when executed by the computer system, cause the computer system to:
 display, in visual association with a first projection of the first span in the natural language text and a second projection of the second span in the natural language text, the information object attributes associated with the first span and the second span; and   accept user input to perform at least one of: confirming the information object attributes or modifying the information object attributes.   
     
     
         20 . The computer-readable non-transitory storage medium of  claim 17 , further comprising executable instructions that, when executed by the computer system, cause the computer system to:
 determine, using a training data set, a parameter of the first classifier model, wherein the training data set comprises an annotated natural language text comprising a plurality of textual annotations, wherein each textual annotation is associated with an information object attribute of an information object of a known category.

Join the waitlist — get patent alerts

Track US2020057810A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.