US2023021040A1PendingUtilityA1
Methods and systems for automated table detection within documents
Est. expiryDec 4, 2038(~12.4 yrs left)· nominal 20-yr term from priority
G06V 30/10G06V 30/412G06V 10/82G06V 30/413G06F 18/214G06F 18/217G06V 30/18057G06V 30/414G06F 18/23G06K 9/6262G06K 9/6256G06K 9/6218
67
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Methods and systems for detecting tables within documents are provided. The methods and systems may include receiving a text of the document that includes a plurality of words depicted in the document image. Feature sets may be calculated for the words and may contain one or more features of a corresponding word of the text. Candidate table words may then be identified based on the features vectors, and may then be used to identify a table location within the document image. In some cases, the candidate table words may be identified using a machine learning model.
Claims
exact text as granted — not AI-modified1 . A method comprising:
receiving a text of a document image that includes a plurality of words depicted in the document image; calculating a plurality of feature sets for the plurality of words, wherein each feature set contains information indicative of one or more features of a corresponding word of the plurality of words; identifying candidate table words among the plurality of words based on the feature sets; identifying, with a clustering procedure, a cluster of candidate table words that correspond to a table within the document image; and defining a candidate table location including a candidate table border of the table that contains the cluster of candidate table words.
2 . The method of claim 1 , wherein the features of the feature set include one or more text features selected from the group consisting of: orthographic properties of the corresponding word, syntactic properties of the corresponding word, and formatting properties of the corresponding word.
3 . The method of claim 1 , wherein the features of the feature set include one or more spatial features selected from the group consisting of: a nearby ruler line distance, a neighbor alignment measurement, and a neighbor distance measurement.
4 . The method of claim 1 , wherein the candidate table border is defined as the rectangle with the smallest area that contains the cluster of candidate table words.
5 . The method of claim 1 , wherein the clustering procedure is a density-based spatial clustering of applications with noise (DBSCAN) procedure.
6 . The method of claim 1 , wherein the candidate table border contains one or more words of the text that are not candidate table words.
7 . The method of claim 1 , further comprising:
predicting a reading order for at least a subset of the words of the text.
8 . The method of claim 7 , wherein the text further includes a location for the subset of words, and wherein predicting the reading order further comprises:
assigning a first word of the subset of words as coming before a second word of the subset of words in the reading order if one or more of the following conditions are true: (i) the second word is below the first word according to the location of the first and second words, or (ii) the second word is at the same height as the first word and is positioned to the right of the first word according to the location.
9 . The method of claim 1 , wherein the candidate table words are identified using a machine learning model.
10 . The method of claim 9 , wherein the machine learning model is a recurrent neural network or a convolutional neural network.
11 . The method of claim 9 , further comprising:
receiving a training text of a training document, including a plurality of words depicted in the training document and a labeled document image indicating a labeled table location of a table within the training document; calculating a plurality of training feature sets for the words of the training text, wherein each training feature set contains information indicative of one or more features of a corresponding word of the plurality of words of the training text; identifying, with the machine learning model, candidate training table words of the training text among the words of the training text based on the training feature sets; identifying, with the clustering procedure, a cluster of candidate training table words that correspond to a table within the document image; defining a candidate training table location including a candidate training table border of the training table that contains the cluster of candidate training table words; comparing the training table location with the labeled table location to identify a table location error of the training table location; and updating one or more parameters of the machine learning model based on the table location error.
12 . A system comprising:
a processor; and a memory containing instructions that, when executed by the processor, cause the processor to:
receive a text of a document image that includes a plurality of words depicted in the document image;
calculate a plurality of feature sets for the words, wherein each feature set contains information indicative of one or more features of a corresponding word of the plurality of words;
identify candidate table words among the plurality of words based on the feature sets;
identify, with a clustering procedure, a cluster of candidate table words that correspond to a table within the document image; and
define a candidate table location including a candidate table border of the table that includes the cluster of candidate table words.
13 . The system of claim 12 , wherein the features of the feature set include one or more text features selected from the group consisting of: orthographic properties of the corresponding word, syntactic properties of the corresponding word, and formatting properties of the corresponding word.
14 . The system of claim 12 , wherein the features of the feature set include one or more spatial features selected from the group consisting of: a nearby ruler line distance, a neighbor alignment measurement, and a neighbor distance measurement.
15 . The system of claim 12 , wherein the candidate table border is defined as the rectangle with the smallest area that contains the cluster of candidate table words.
16 . The system of claim 12 , wherein the memory contains further instructions which, when executed by the processor, cause the processor to:
predict a reading order for at least a subset of the words of the text.
17 . The system of claim 16 , wherein the text further includes a location for the subset of words, and wherein the memory contains further instructions which, when executed by the processor, cause the processor to:
assign a first word of the subset of words as coming before a second word of the subset of words in the reading order if one or more of the following conditions are true: (i) the second word is below the first word according to the location of the first and second words, or (ii) the second word is at the same height as the first word and is positioned to the right of the first word according to the location.
18 . The system of claim 12 , wherein the text classifier includes a machine learning model configured to identify the candidate table words, the machine learning model including at least one of a recurrent neural network and a convolutional neural network.
19 . The system of claim 18 , further comprising a training system configured, when executed by the processor, to:
receive a training text of a training document, including a plurality of words depicted in the training document and a labeled document image indicating a labeled table location of a table within the training document; calculate a plurality of training feature sets for the words of the training text, wherein each training feature set contains one or more features of a corresponding word of the plurality of words of the training text; identify, with the machine learning model, candidate training table words of the training text among the words of the training text; identify, with the clustering procedure, a cluster of candidate training table words that correspond to a table within the document image; define a candidate training table location including a candidate training table border of the training table that contains the cluster of candidate training table words; compare the training table location with the labeled table location to identify a table location error of the training table location; and update one or more parameters of the machine learning model based on the table location error.
20 . A computer-readable medium containing instructions which, when executed by a processor, cause the processor to:
receive a text of a document image that includes a plurality of words depicted in the document image; calculate a plurality of feature sets for the words, wherein each feature set contains information indicative of one or more features of a corresponding word of the plurality of words; identify candidate table words among the plurality of words based on the feature sets; identify, with a clustering procedure, a cluster of candidate table words that correspond to a table within the document image; and define a candidate table location including a candidate table border of the table that includes the cluster of candidate table words.Join the waitlist — get patent alerts
Track US2023021040A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.