Optimization techniques for artificial intelligence
Abstract
Methods, apparatuses and computer readable medium are presented for generating a natural language model. A method for generating a natural language model comprises: selecting from a pool of documents, a first set of documents to be annotated; receiving annotations of the first set of documents elicited by first human readable prompts; training a natural language model using the annotated first set of documents; determining documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations; selecting from the pool of documents, a second set of documents to be annotated comprising documents having uncertain natural language processing results; receiving annotations of the second set of documents elicited by second human readable prompts; and retraining a natural language model using the annotated second set of documents.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for generating a natural language model, the method comprising:
selecting by one or more processors in a natural language platform, from a pool of documents, a first set of documents to be annotated; for each document in the first set of documents, generating, by the one or more processors, a first human readable prompt configured to elicit an annotation of said document; receiving annotations of the first set of documents elicited by the first human readable prompts; training, by the one or more processors, a natural language model using the annotated first set of documents; determining, by the one or more processors, documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations; selecting by the one or more processors, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results; for each document in the second set of documents, generating, by the one or more processors, a second human readable prompt configured to elicit an annotation of said document; receiving annotations of the second set of documents elicited by the second human readable prompts; and retraining, by the one or more processors, a natural language model using the annotated second set of documents.
2 . The method of claim 1 , wherein the annotations of the first and second sets of documents comprise classification of the documents into one or more categories among a plurality of categories.
3 . The method of claim 1 , wherein the annotations of the first and second sets of documents comprise selection of one or more portions of the documents relevant to one or more topics.
4 . The method of claim 1 , wherein the steps of determining documents having uncertain natural language processing results; selecting a second set of documents to be annotated; generating a second human readable prompt; receiving annotations of the second set of documents; and retraining a natural language model using the annotated second set of documents, are repeated until the trained model has reached a predetermined performance level.
5 . The method of claim 1 , wherein selecting the first set of documents comprises selecting documents that are evenly distributed among different document types.
6 . The method of claim 1 , wherein selecting the first set of documents comprises selecting at least one document within each of a plurality of machine-discovered topics.
7 . The method of claim 1 , wherein selecting the first set of documents comprises selecting documents based on a keyword search.
8 . The method of claim 1 , wherein selecting the first set of documents comprises selecting documents based on confidence levels generated by analysis of the documents by one or more pre-existing natural language models.
9 . The method of claim 1 , wherein selecting the first set of documents comprises a manual selection.
10 . The method of claim 1 , wherein selecting the first set of documents comprises removing exact duplicates and/or near duplicates from the first set of documents.
11 . The method of claim 1 , wherein selecting the first set of documents comprises selecting documents based on features contained therein.
12 . The method of claim 1 , wherein determining documents having uncertain natural language processing results is based on confidence levels generated by analysis of the documents by the trained model.
13 . The method of claim 1 , further comprising training, by the one or more processors, a plurality of additional natural language models using the annotated first set of documents, wherein determining documents having uncertain natural language processing results is based on a level of disagreement among the plurality of additional natural language models.
14 . The method of claim 13 , wherein the level of disagreement is determined by assigning more weight to models with better known performance levels than models with worse known performance levels.
15 . The method of claim 1 , wherein determining documents having uncertain natural language processing results is based on a level of disagreement among more than one annotator.
16 . The method of claim 15 , wherein the level of disagreement is determined by assigning more weight to annotators with better known performance levels than annotators with worse known performance levels.
17 . The method of claim 15 , wherein selecting the second set of documents comprises selecting documents similar to documents that have a high level of disagreement among more than one annotator.
18 . The method of claim 1 , wherein the second human readable prompt is configured to elicit a true-or-false answer aimed at resolving uncertainty in the natural language processing results.
19 . An apparatus for generating a natural language model, the apparatus comprising one or more processors configured to:
select, from a pool of documents, a first set of documents to be annotated; for each document in the first set of documents, generate a first human readable prompt configured to elicit an annotation of said document; receive annotations of the first set of documents elicited by the first human readable prompts; train a natural language model using the annotated first set of documents; determine documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations; select, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results; for each document in the second set of documents, generate a second human readable prompt configured to elicit an annotation of said document; receive annotations of the second set of documents elicited by the second human readable prompts; and retrain a natural language model using the annotated second set of documents.
20 . A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to:
select, from a pool of documents, a first set of documents to be annotated; for each document in the first set of documents, generate a first human readable prompt configured to receive annotations of the first set of documents; train a natural language model using the annotated first set of documents; determine documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations; select, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results; for each document in the second set of documents, generate a second human readable prompt configured to receive annotations of the second set of documents; and retrain a natural language model using the annotated second set of documents.Join the waitlist — get patent alerts
Track US2020234002A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.