US2021342718A1PendingUtilityA1

Method for training information retrieval model based on weak-supervision and method for providing search result using such model

Assignee: UNIV HOSEO ACAD COOP FOUNDPriority: May 1, 2020Filed: Apr 30, 2021Published: Nov 4, 2021
Est. expiryMay 1, 2040(~13.8 yrs left)· nominal 20-yr term from priority
G06N 20/00G06F 16/334G06F 16/3329G06F 16/3334G06F 40/258G06F 16/93G06N 5/04G06F 16/24578
55
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A search method using an artificial intelligence based information retrieval model and a method for training the artificial intelligence based information retrieval model used for the method are provided.In the method, even if there is no labeled data and only a corpus exists, the artificial intelligence based information retrieval model can be trained using the weak-supervision methodology. Search can be performed by dividing documents into passages having short lengths. Compared to an information retrieval model based on unsupervised learning, improved search results are provided.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computer-implemented method for providing a user with a search result corresponding to a query input by the user from a corpus including a plurality of passages, comprising the steps of:
 (a) retrieving, by an information retrieval model based on an unsupervised learning, N passages corresponding to the query from the corpus;   (b) re-ranking, by an artificial intelligence based information retrieval model trained by a weak-supervision methodology, the N passages retrieved in step (a) based on the query; and,   (c) outputting a list of search results corresponding to the passages re-ranked in step (b),   wherein each passage of the corpus is a part of each document of the document corpus including a plurality of documents to be searched, and   wherein the artificial intelligence based information retrieval model trained by the weak-supervision methodology is trained based on pseudo-labels which are generated by using the plurality of passages included in the corpus and a pseudo-query generated from each passage, said pseudo-query generated by an information retrieval model based on unsupervised learning including the information retrieval model that is used in the step (a).   
     
     
         2 . The method according to  claim 1 , wherein each document in the document corpus has a title, and wherein each passage in the corpus includes the title of the document in which the passage is included as part of. 
     
     
         3 . The method according to  claim 1 , wherein, in the step (c), the search result list is a list in which documents including the re-ranked passages are sorted so as to correspond to the sorting order of the re-ranked passages in step (b). 
     
     
         4 . The method according to  claim 3 , wherein, when the re-ranked result in step (b) includes more than two passages extracted from one document, the sorting order of the document is determined to correspond to the order of the most relevant passage. 
     
     
         5 . An apparatus for providing a user with a search result corresponding to a query input by the user from a corpus including a plurality of passages, comprising:
 at least one processor; and   at least one memory for storing computer-executable instructions,   the computer-executable instructions stored in the at least one memory that are operable, when executed by the at least one processor, to perform operations including:   (a) retrieving, by an information retrieval model based on an unsupervised learning, N passages corresponding to the query from the corpus;   (b) re-ranking, by an artificial intelligence based information retrieval model trained by a weak-supervision methodology, the N passages retrieved in step (a) based on the query; and,   (c) outputting a list of search results corresponding to the passages re-ranked in step (b),   wherein each passage of the corpus is a part of each document of the document corpus including a plurality of documents to be searched, and   wherein the artificial intelligence based information retrieval model trained by the weak-supervision methodology is trained based on pseudo-labels which are generated by using the plurality of passages included in the corpus and a pseudo-query generated from each passage, said pseudo-query generated by an information retrieval model based on unsupervised learning including the information retrieval model that is used in the step (a).   
     
     
         6 . The apparatus according to  claim 5 , wherein each document in the document corpus has a title, and wherein each passage in the corpus includes the title of the document in which the passage is included as part of. 
     
     
         7 . The apparatus according to  claim 5 , wherein, in the step (c), the search result list is a list in which documents including the re-ranked passages are sorted so as to correspond to the sorting order of the re-ranked passages in step (b). 
     
     
         8 . The apparatus according to  claim 7 , wherein, when the re-ranked result in step (b) includes more than two passages extracted from one document, the sorting order of the document is determined to correspond to the order of the most relevant passage.

Join the waitlist — get patent alerts

Track US2021342718A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.