US2024046159A1PendingUtilityA1

Continual learning for multi modal systems using crowd sourcing

Assignee: CISCO TECH INCPriority: Apr 13, 2018Filed: Oct 19, 2023Published: Feb 8, 2024
Est. expiryApr 13, 2038(~11.7 yrs left)· nominal 20-yr term from priority
G06N 20/00G10L 15/06G10L 15/063G10L 15/02G06F 18/23G06F 18/41G06F 18/253G06V 10/987G06V 10/774G10L 2015/0631
74
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems, methods, and devices are disclosed for training a model. Media data is separated into one or more clusters, each cluster based on a feature from a first model. The media data of each cluster is sampled and, based on an analysis of the sampled media data, an accuracy of the media data of each cluster is determined. The accuracy is associated with the feature from the first model. Based on a subset dataset of the media data being outside a threshold accuracy, the subset dataset is automatically forwarded to a crowd source service. Verification of the subset dataset is received from the crowd source service, and the verified subset dataset is added to the first model.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A system comprising:
 at least one processor and at least one memory containing instructions that, when executed, cause the at least one processor to:
 receive a continuous pipeline of crowd sourced data; 
 separate the data into one or more clusters, each cluster based at least on a feature from one or more models; 
 determine an accuracy associated with each respective feature of a subset dataset of the data of each cluster; 
 automatically detect, based at least in part on the accuracy, that the subset dataset of the data is outside of a threshold accuracy; 
 in response to automatically detecting that the accuracy being outside of the threshold accuracy, automatically forward the subset dataset to a crowd source service; 
 receive verification of the subset dataset from the crowd source service; and 
 add the verified subset dataset to at least one model of the one or more models. 
   
     
     
         2 . The system of  claim 1 , the at least one processor further configured to:
 generate a second model based on the received verified subset dataset; and   upon determining that an accuracy of the second model exceeds the at least one model, update the at least one model with the second model.   
     
     
         3 . The system of  claim 1 , wherein the at least one processor that determines the accuracy of the subset dataset is configured to:
 determine a feature metric associated with the respective feature;   define a centroid based on the feature metric; and   determine a measured metric associated with the data in each cluster.   
     
     
         4 . The system of  claim 3 , wherein the at least one processor is further configured to:
 select the subset dataset of each cluster based on the measured metric matching the feature metric associated with the centroid within a threshold.   
     
     
         5 . The system of  claim 1 , wherein the at least one processor is further configured to:
 receive a labelled subset dataset from the crowd source service;   add the labelled subset dataset to the least one model to create a combined model dataset;   generate a second model based on the combined model dataset; and   determine the accuracy of the second model.   
     
     
         6 . The system of  claim 5 , wherein the at least one processor is configured to generate the second model on an ongoing basis. 
     
     
         7 . The system of  claim 1 , wherein the at least one processor is configured to determine the feature is at least one of an accent, gender, or environmental background noise in the data. 
     
     
         8 . The system of  claim 1 , wherein the data is comprised of multiple data types, the multiple data types including audio, visual, and text data. 
     
     
         9 . The system of  claim 1 , wherein the data is configured to be separated into the one or more clusters based on an unsupervised machine learning technique. 
     
     
         10 . The system of  claim 1 , wherein the at least one processor is further configured to automatically forward the subset dataset to the crowd source service based on a volume of the subset dataset being outside the threshold accuracy. 
     
     
         11 . The system of  claim 10 , wherein the volume of the subset dataset that initiates forwarding to the crowd source service is based on a volume heuristics model that is configured to determine an amount of data predicted to successfully update the at least one model. 
     
     
         12 . A non-transitory computer-readable medium comprising instructions, when executed by at least one processor of a system, causes the system to:
 receive a continuous pipeline of crowd sourced data;   separate the data into one or more clusters, each cluster based at least on a feature from one or more models;   determine an accuracy associated with each respective feature of a subset dataset of the data of each cluster;   automatically detect, based at least in part on the accuracy, that the subset dataset of the data is outside of a threshold accuracy;   in response to automatically detecting that the accuracy being outside of the threshold accuracy, automatically forward the subset dataset to a crowd source service;   receive verification of the subset dataset from the crowd source service; and   add the verified subset dataset to at least one model of the one or more models.   
     
     
         13 . The non-transitory computer-readable medium of  claim 12 , wherein the at least one processor is further configured to:
 generate a second model based on the received verified subset dataset; and upon determining that an accuracy of the second model exceeds the at least one model, update the at least one model with the second model.   
     
     
         14 . The non-transitory computer-readable medium of  claim 12 , wherein the at least one processor is further configured to:
 determine a feature metric associated with the respective feature;   define a centroid based on the feature metric; and   determine a measured metric associated with the data in each cluster.   
     
     
         15 . The non-transitory computer-readable medium of  claim 12 , wherein the at least one processor is further configured to:
 receive a labelled subset dataset from the crowd source service;   add the labelled subset dataset to the least one model to create a combined model dataset;   generate a second model based on the combined model dataset; and   determine the accuracy of the second model.   
     
     
         16 . A method comprising:
 receiving a continuous pipeline of crowd sourced data;   separating the data into one or more clusters, each cluster based at least on a feature from one or more models;   determining an accuracy associated with each respective feature of a subset dataset of the data of each cluster;   automatically detecting, based at least in part on the accuracy, that the subset dataset of the data is outside of a threshold accuracy;   in response to automatically detecting that the accuracy being outside of the threshold accuracy, automatically forwarding the subset dataset to a crowd source service;   receiving verification of the subset dataset from the crowd source service; and   adding the verified subset dataset to at least one model of the one or more models.   
     
     
         17 . The method of  claim 16 , further comprising:
 generating a second model based on the received verified subset dataset; and   upon determining that an accuracy of the second model exceeds the at least one model, update the at least one model with the second model.   
     
     
         18 . The method of  claim 16 , further comprising:
 receiving a labelled subset dataset from the crowd source service;   adding the labelled subset dataset to the least one model to create a combined model dataset;   generating a second model based on the combined model dataset; and   determining the accuracy of the second model.   
     
     
         19 . The method of  claim 18 , wherein generating the second model is on an ongoing basis. 
     
     
         20 . The method of  claim 16 , wherein the data is configured to be separated into the one or more clusters based on an unsupervised machine learning technique.

Join the waitlist — get patent alerts

Track US2024046159A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.