System and method to extract information from unstructured image documents
Abstract
The present disclosure relates to a system and method to extract information from unstructured image documents. The extraction technique is content-driven and not dependent on the layout of a particular image document type. The disclosed method breaks down an image document into smaller images using the text cluster detection algorithm. The smaller images are converted into text samples using optical character recognition (OCR). Each of the text samples is fed to a trained machine learning model. The model classifies each text sample into one of a plurality of pre-determined field types. The desired value extraction problem may be converted into a question-answering problem using a pre-trained model. A fixed question is formed on the basis of the classified field type. The output of the question-answering model may be passed through a rule-based post-processing step to obtain the final answer.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system for recognizing a relevant value from an unstructured document, where a computer processor performs the operations of:
receiving an unstructured document as input; detecting a plurality of text clusters in the unstructured document; generating, by an optical character recognition (OCR) module, a plurality of text outputs from the plurality of text clusters, wherein each text cluster corresponds to a respective text output; classifying the plurality of text outputs using a natural language processing algorithm configured to classify text; using a pre-trained question-answering model to obtain an initial answer from one or more of the classified plurality of text outputs; and extract a final answer, based on the initial answer, to be presented as an extracted value to be associated with a corresponding field.
2 . The system of claim 1 , wherein detecting the text clusters is agnostic of a layout or format of the unstructured document.
3 . The system of claim 1 , wherein the unstructured document in an unstructured image document in its entirety, or the unstructured document has a portion that is an unstructured image document.
4 . The system of claim 3 , wherein each of the text clusters is a smaller image within the unstructured image document.
5 . The system of claim 1 , wherein each of the text clusters is bounded within a contour of a respective bounding box.
6 . The system of claim 5 , where two neighboring bounding boxes are merged based on proximity of individual neighboring bounding box coordinates.
7 . The system of claim 5 , wherein the computer processor further performs the operation of:
checking if a desired value is extracted from an initial set of text clusters with their respective bounding boxes; responsive to determining that the desired value is not extracted from the initial set of text clusters with their respective bounding boxes, creating compound bounding boxes by merging one or more neighboring bounding boxes, thereby merging the corresponding text clusters.
8 . The system of claim 7 , wherein the computer processor further performs the operation of:
generating a new text output by performing an optical character recognition (OCR) operation on the merged text clusters; classifying the new text output using the natural language processing algorithm configured to classify text; using the pre-trained question-answering model to obtain a revised answer from the classified new text output; and extract a revised final answer, based on the revised initial answer, to be presented as a new extracted value to be associated with the field.
9 . The system of claim 1 , wherein detecting the plurality of text clusters uses a text cluster detection algorithm executed by the computer processor.
10 . The system of claim 9 , wherein the text cluster detection algorithm applies a morphological transformation on the bounding boxes along one or more axes.
11 . The system of claim 10 , wherein the morphological transformation is an iterative process that, in each iteration, creates an intermediate set of bounding boxes from a previous set of bounding boxes.
12 . The system of claim 11 , wherein in one or more iteration, the morphological transformation along a selected axis applies a dilation rate that is faster or slower than a dilation rate along another axis.
13 . The system of claim 11 , wherein relative dilation rates along different axes can be scaled up or down by predetermined factors.
14 . The system of claim 11 , wherein dilation scaling can be applied to axes of the bounding boxes sequentially.
15 . The system of claim 11 , wherein dilation scaling can be applied to both axes of the bounding boxes in parallel.
16 . The system of claim 11 , wherein the intermediate set of bounding boxes are fed to the OCR module to obtain an array of text samples for natural language processing.
17 . The method of claim 1 , where the pre-trained question answering model is used to fetch one or more relevant answers corresponding to each of the plurality of text outputs.
18 . The method of claim 17 , where an output of the question-answering model is passed through one or more rule-based filters to obtain the final answer.Join the waitlist — get patent alerts
Track US2024013563A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.