US2020234002A1PendingUtilityA1

Optimization techniques for artificial intelligence

Assignee: AIPARC HOLDINGS PTE LTDPriority: Dec 9, 2014Filed: Nov 21, 2018Published: Jul 23, 2020
Est. expiryDec 9, 2034(~8.4 yrs left)· nominal 20-yr term from priority
G06Q 10/40G06F 40/169G06F 40/42G06F 40/30G06F 16/3329G06F 3/0482G06F 16/24532G06F 16/285G06F 16/243G06F 16/367G06F 40/40G06F 16/288G06N 20/00G06F 16/35G06F 16/93G06F 40/221G06F 16/951G06F 40/137G06Q 50/01
71
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods, apparatuses and computer readable medium are presented for generating a natural language model. A method for generating a natural language model comprises: selecting from a pool of documents, a first set of documents to be annotated; receiving annotations of the first set of documents elicited by first human readable prompts; training a natural language model using the annotated first set of documents; determining documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations; selecting from the pool of documents, a second set of documents to be annotated comprising documents having uncertain natural language processing results; receiving annotations of the second set of documents elicited by second human readable prompts; and retraining a natural language model using the annotated second set of documents.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for generating a natural language model, the method comprising:
 selecting by one or more processors in a natural language platform, from a pool of documents, a first set of documents to be annotated;   for each document in the first set of documents, generating, by the one or more processors, a first human readable prompt configured to elicit an annotation of said document;   receiving annotations of the first set of documents elicited by the first human readable prompts;   training, by the one or more processors, a natural language model using the annotated first set of documents;   determining, by the one or more processors, documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations;   selecting by the one or more processors, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results;   for each document in the second set of documents, generating, by the one or more processors, a second human readable prompt configured to elicit an annotation of said document;   receiving annotations of the second set of documents elicited by the second human readable prompts; and   retraining, by the one or more processors, a natural language model using the annotated second set of documents.   
     
     
         2 . The method of  claim 1 , wherein the annotations of the first and second sets of documents comprise classification of the documents into one or more categories among a plurality of categories. 
     
     
         3 . The method of  claim 1 , wherein the annotations of the first and second sets of documents comprise selection of one or more portions of the documents relevant to one or more topics. 
     
     
         4 . The method of  claim 1 , wherein the steps of determining documents having uncertain natural language processing results; selecting a second set of documents to be annotated; generating a second human readable prompt; receiving annotations of the second set of documents; and retraining a natural language model using the annotated second set of documents, are repeated until the trained model has reached a predetermined performance level. 
     
     
         5 . The method of  claim 1 , wherein selecting the first set of documents comprises selecting documents that are evenly distributed among different document types. 
     
     
         6 . The method of  claim 1 , wherein selecting the first set of documents comprises selecting at least one document within each of a plurality of machine-discovered topics. 
     
     
         7 . The method of  claim 1 , wherein selecting the first set of documents comprises selecting documents based on a keyword search. 
     
     
         8 . The method of  claim 1 , wherein selecting the first set of documents comprises selecting documents based on confidence levels generated by analysis of the documents by one or more pre-existing natural language models. 
     
     
         9 . The method of  claim 1 , wherein selecting the first set of documents comprises a manual selection. 
     
     
         10 . The method of  claim 1 , wherein selecting the first set of documents comprises removing exact duplicates and/or near duplicates from the first set of documents. 
     
     
         11 . The method of  claim 1 , wherein selecting the first set of documents comprises selecting documents based on features contained therein. 
     
     
         12 . The method of  claim 1 , wherein determining documents having uncertain natural language processing results is based on confidence levels generated by analysis of the documents by the trained model. 
     
     
         13 . The method of  claim 1 , further comprising training, by the one or more processors, a plurality of additional natural language models using the annotated first set of documents, wherein determining documents having uncertain natural language processing results is based on a level of disagreement among the plurality of additional natural language models. 
     
     
         14 . The method of  claim 13 , wherein the level of disagreement is determined by assigning more weight to models with better known performance levels than models with worse known performance levels. 
     
     
         15 . The method of  claim 1 , wherein determining documents having uncertain natural language processing results is based on a level of disagreement among more than one annotator. 
     
     
         16 . The method of  claim 15 , wherein the level of disagreement is determined by assigning more weight to annotators with better known performance levels than annotators with worse known performance levels. 
     
     
         17 . The method of  claim 15 , wherein selecting the second set of documents comprises selecting documents similar to documents that have a high level of disagreement among more than one annotator. 
     
     
         18 . The method of  claim 1 , wherein the second human readable prompt is configured to elicit a true-or-false answer aimed at resolving uncertainty in the natural language processing results. 
     
     
         19 . An apparatus for generating a natural language model, the apparatus comprising one or more processors configured to:
 select, from a pool of documents, a first set of documents to be annotated;   for each document in the first set of documents, generate a first human readable prompt configured to elicit an annotation of said document;   receive annotations of the first set of documents elicited by the first human readable prompts;   train a natural language model using the annotated first set of documents;   determine documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations;   select, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results;   for each document in the second set of documents, generate a second human readable prompt configured to elicit an annotation of said document;   receive annotations of the second set of documents elicited by the second human readable prompts; and   retrain a natural language model using the annotated second set of documents.   
     
     
         20 . A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to:
 select, from a pool of documents, a first set of documents to be annotated;   for each document in the first set of documents, generate a first human readable prompt configured to receive annotations of the first set of documents;   train a natural language model using the annotated first set of documents;   determine documents in the pool having uncertain natural language processing results according to the trained natural language model and/or the received annotations;   select, from the pool of documents, a second set of documents to be annotated comprising one or more of the documents having uncertain natural language processing results;   for each document in the second set of documents, generate a second human readable prompt configured to receive annotations of the second set of documents; and   retrain a natural language model using the annotated second set of documents.

Join the waitlist — get patent alerts

Track US2020234002A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.