Method and apparatus for training document information extraction model, and method and apparatus for extracting document information
Abstract
The present disclosure provides a method and apparatus for training a document information extraction model and method and apparatus for extracting document information, and relates to the field of artificial intelligence, and more particularly to the field of natural language processing. A specific implementation solution is: acquiring training data labeled with an answer corresponding to a preset question and a document information extraction model, the training data includes layout document training data and streaming document training data; extracting at least one feature from the training data; fusing at least one feature to obtain a fused feature; inputting the preset question, the fused feature and the training data into the document information extraction model to obtain a predicted result; and adjusting network parameters of the document information extraction model based on the predicted result and the answer.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for training a document information extraction model, comprising:
acquiring training data labeled with an answer corresponding to a preset question and a document information extraction model, wherein the training data comprises layout document training data and streaming document training data; extracting at least one feature from the training data; fusing the at least one feature to obtain a fused feature; inputting the preset question, the fused feature, and the training data into the document information extraction model to obtain a predicted result; and adjusting network parameters of the document information extraction model based on the predicted result and the answer.
2 . The method of claim 1 , wherein acquiring training data labeled with an answer corresponding to a preset question, comprises:
acquiring text content of a web page and corresponding key-value pair information by crawling and parsing the web page; and constructing a streaming document training data labeled with the answer corresponding to the preset question according to the text content and the corresponding key-value pair information.
3 . The method of claim 1 , wherein acquiring training data labeled with an answer corresponding to a preset question, comprises:
acquiring the streaming document training data and a layout document set; emptying text content in the layout document set, and retaining a document structure; and filling the streaming document training data into the document structure to generate the layout document training data.
4 . The method of claim 1 , wherein extracting at least one feature from the training data, comprises:
extracting at least one of streaming reading order information, spatial position information of text characters, text segmentation information or a document type from the training data.
5 . A method for extracting document information, comprising:
acquiring document information to be extracted; extracting at least one feature from the document information; fusing the at least one feature to obtain a fused feature; inputting a preset question, the fused feature, and the document information into a document information extraction model trained by a method for training the document information extraction model to obtain an answer, the method for training a document information extraction model comprising: acquiring training data labeled with an answer corresponding to the preset question and the document information extraction model, wherein the training data comprises layout document training data and streaming document training data; extracting at least one feature from the training data; fusing the at least one feature to obtain a fused feature; inputting the preset question, the fused feature, and the training data into the document information extraction model to obtain a predicted result; and adjusting network parameters of the document information extraction model based on the predicted result and the answer.
6 . An electronic device, comprising:
at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor to cause the at least one processor to perform operations for training a document information extraction model, the operations comprising: acquiring training data labeled with an answer corresponding to a preset question and a document information extraction model, wherein the training data comprises layout document training data and streaming document training data; extracting at least one feature from the training data; fusing the at least one feature to obtain a fused feature; inputting the preset question, the fused feature, and the training data into the document information extraction model to obtain a predicted result; and adjusting network parameters of the document information extraction model based on the predicted result and the answer.
7 . The electronic device of claim 6 , wherein acquiring training data labeled with an answer corresponding to a preset question, comprises:
acquiring text content of a web page and corresponding key-value pair information by crawling and parsing the web page; and constructing a streaming document training data labeled with the answer corresponding to the preset question according to the text content and the corresponding key-value pair information.
8 . The electronic device of claim 6 , wherein acquiring training data labeled with an answer corresponding to a preset question, comprises:
acquiring the streaming document training data and a layout document set; emptying text content in the layout document set, and retaining a document structure; and filling the streaming document training data into the document structure to generate the layout document training data.
9 . The electronic device of claim 6 , wherein extracting at least one feature from the training data, comprises:
extracting at least one of streaming reading order information, spatial position information of text characters, text segmentation information or a document type from the training data.Join the waitlist — get patent alerts
Track US2023177359A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.