US2020089675A1PendingUtilityA1

Methods and apparatuses for iterative data mining

Assignee: FRAUNHOFER GES FORSCHUNGPriority: Apr 4, 2014Filed: Nov 22, 2019Published: Mar 19, 2020
Est. expiryApr 4, 2034(~7.7 yrs left)· nominal 20-yr term from priority
G06F 16/2465
48
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

One or more data mining processes may be executed based on control parameters to discover a plurality of result patterns in a data set. The discovered result patterns are presented to a user. Information on one or more selected result patterns, where the selection involves the user's subjective interest, is received. The control parameters are automatically updated based on the received information on the selected result patterns.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for mining data from a data set, the method comprising:
 performing a pattern discovery round based on a plurality of control parameters, the pattern discovery round comprising:
 executing one or more of a plurality of data mining processes based on the control parameters to discover a plurality of result patterns in a data set; 
 presenting the discovered result patterns to a user; 
 receiving information on one or more result patterns selected by the user; 
 automatically updating the plurality control parameters based on the selected result patterns; and 
   performing a subsequent pattern discovery round based on the updated control parameters.   
     
     
         2 . The method of  claim 1 , wherein the plurality of control parameters are computed based on:
 a pattern utility model, and   a process selection probability distribution.   
     
     
         3 . The method of  claim 2 , wherein the pattern utility model is a linear model of the form
     u   t ( p )=   w   t ,φ( p ) ,
   where φ: P→   d  denotes a d-dimensional feature map from the space of possible patterns p, and w t  ∈   d  denotes a model parameter approximation at time t.   
     
     
         4 . The method of  claim 1 , wherein executing one or more of a plurality of data mining processes comprises executing the data mining processes on a programmable hardware device in parallel or serially. 
     
     
         5 . The method of  claim 1 , wherein presenting the discovered result patterns to the user comprises displaying the result patterns to the user via a graphical user interface. 
     
     
         6 . The method of  claim 1 , wherein receiving information on one or more result patterns comprises receiving information on a relevance of the one or more result patterns. 
     
     
         7 . The method of  claim 1 , wherein prior to executing one or more of a plurality of data mining processes the method comprises randomly selecting the data mining process from the plurality of data mining processes based on a process selection probability distribution function. 
     
     
         8 . The method of  claim 7 , wherein the process selection probability distribution function for a pattern discovery round iteration 1 and data mining algorithm i is computed according to
   π l,i =((γ l −1) v   i )/ V+γ   l   /k,  
   wherein V is a normalization factor, v i  is a vector of performance potential weights for algorithm i, k denotes the total number of data mining algorithms, and γ l  denotes a bandit mixture coefficient.   
     
     
         9 . The method of  claim 1 , further comprising:
 upon executing a data mining process, adding one or more discovered result patterns to a pattern cache memory, leading to a change of state of the pattern cache memory.   
     
     
         10 . The method of  claim 9 , further comprising:
 assessing a performance of the executed data mining process based on the change of state of the pattern cache memory and a current pattern utility model.   
     
     
         11 . The method of  claim 10 , wherein a process selection probability distribution is updated based on the assessed performance and wherein a next data mining process from the plurality of data mining processes is randomly selected for execution based on the updated process selection probability distribution. 
     
     
         12 . The method of  claim 1 , wherein presenting the discovered result patterns comprises proposing a ranking of the at least one result pattern based on a current state of a pattern cache memory storing one or more discovered result patterns. 
     
     
         13 . The method of  claim 12 , wherein the proposed ranking of candidate patterns is computed based on a current pattern utility model. 
     
     
         14 . The method of  claim 12 , wherein the proposed ranking of candidate patterns is computed based on a greedy algorithm process that maximizes a ranking utility function at each stage. 
     
     
         15 . The method of  claim 12  further comprising determining a feedback ranking of patterns based on a pattern utility model in one or more of the result patterns of the proposed ranking of candidate patterns. 
     
     
         16 . The method of  claim 15 , wherein the feedback ranking is determined based on candidate patterns that have been declared by the user as relevant and based on candidate patterns that have been declared by the user as irrelevant. 
     
     
         17 . The method of  claim 16 , wherein updating the pattern utility model is based on a comparison of the feedback ranking with the proposed ranking of candidate patterns.

Join the waitlist — get patent alerts

Track US2020089675A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.