Continual learning for multi modal systems using crowd sourcing
Abstract
Systems, methods, and devices are disclosed for training a model. Media data is separated into one or more clusters, each cluster based on a feature from a first model. The media data of each cluster is sampled and, based on an analysis of the sampled media data, an accuracy of the media data of each cluster is determined. The accuracy is associated with the feature from the first model. Based on a subset dataset of the media data being outside a threshold accuracy, the subset dataset is automatically forwarded to a crowd source service. Verification of the subset dataset is received from the crowd source service, and the verified subset dataset is added to the first model.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system comprising:
at least one processor and at least one memory containing instructions that, when executed, cause the at least one processor to:
receive a continuous pipeline of crowd sourced data;
separate the data into one or more clusters, each cluster based at least on a feature from one or more models;
determine an accuracy associated with each respective feature of a subset dataset of the data of each cluster;
automatically detect, based at least in part on the accuracy, that the subset dataset of the data is outside of a threshold accuracy;
in response to automatically detecting that the accuracy being outside of the threshold accuracy, automatically forward the subset dataset to a crowd source service;
receive verification of the subset dataset from the crowd source service; and
add the verified subset dataset to at least one model of the one or more models.
2 . The system of claim 1 , the at least one processor further configured to:
generate a second model based on the received verified subset dataset; and upon determining that an accuracy of the second model exceeds the at least one model, update the at least one model with the second model.
3 . The system of claim 1 , wherein the at least one processor that determines the accuracy of the subset dataset is configured to:
determine a feature metric associated with the respective feature; define a centroid based on the feature metric; and determine a measured metric associated with the data in each cluster.
4 . The system of claim 3 , wherein the at least one processor is further configured to:
select the subset dataset of each cluster based on the measured metric matching the feature metric associated with the centroid within a threshold.
5 . The system of claim 1 , wherein the at least one processor is further configured to:
receive a labelled subset dataset from the crowd source service; add the labelled subset dataset to the least one model to create a combined model dataset; generate a second model based on the combined model dataset; and determine the accuracy of the second model.
6 . The system of claim 5 , wherein the at least one processor is configured to generate the second model on an ongoing basis.
7 . The system of claim 1 , wherein the at least one processor is configured to determine the feature is at least one of an accent, gender, or environmental background noise in the data.
8 . The system of claim 1 , wherein the data is comprised of multiple data types, the multiple data types including audio, visual, and text data.
9 . The system of claim 1 , wherein the data is configured to be separated into the one or more clusters based on an unsupervised machine learning technique.
10 . The system of claim 1 , wherein the at least one processor is further configured to automatically forward the subset dataset to the crowd source service based on a volume of the subset dataset being outside the threshold accuracy.
11 . The system of claim 10 , wherein the volume of the subset dataset that initiates forwarding to the crowd source service is based on a volume heuristics model that is configured to determine an amount of data predicted to successfully update the at least one model.
12 . A non-transitory computer-readable medium comprising instructions, when executed by at least one processor of a system, causes the system to:
receive a continuous pipeline of crowd sourced data; separate the data into one or more clusters, each cluster based at least on a feature from one or more models; determine an accuracy associated with each respective feature of a subset dataset of the data of each cluster; automatically detect, based at least in part on the accuracy, that the subset dataset of the data is outside of a threshold accuracy; in response to automatically detecting that the accuracy being outside of the threshold accuracy, automatically forward the subset dataset to a crowd source service; receive verification of the subset dataset from the crowd source service; and add the verified subset dataset to at least one model of the one or more models.
13 . The non-transitory computer-readable medium of claim 12 , wherein the at least one processor is further configured to:
generate a second model based on the received verified subset dataset; and upon determining that an accuracy of the second model exceeds the at least one model, update the at least one model with the second model.
14 . The non-transitory computer-readable medium of claim 12 , wherein the at least one processor is further configured to:
determine a feature metric associated with the respective feature; define a centroid based on the feature metric; and determine a measured metric associated with the data in each cluster.
15 . The non-transitory computer-readable medium of claim 12 , wherein the at least one processor is further configured to:
receive a labelled subset dataset from the crowd source service; add the labelled subset dataset to the least one model to create a combined model dataset; generate a second model based on the combined model dataset; and determine the accuracy of the second model.
16 . A method comprising:
receiving a continuous pipeline of crowd sourced data; separating the data into one or more clusters, each cluster based at least on a feature from one or more models; determining an accuracy associated with each respective feature of a subset dataset of the data of each cluster; automatically detecting, based at least in part on the accuracy, that the subset dataset of the data is outside of a threshold accuracy; in response to automatically detecting that the accuracy being outside of the threshold accuracy, automatically forwarding the subset dataset to a crowd source service; receiving verification of the subset dataset from the crowd source service; and adding the verified subset dataset to at least one model of the one or more models.
17 . The method of claim 16 , further comprising:
generating a second model based on the received verified subset dataset; and upon determining that an accuracy of the second model exceeds the at least one model, update the at least one model with the second model.
18 . The method of claim 16 , further comprising:
receiving a labelled subset dataset from the crowd source service; adding the labelled subset dataset to the least one model to create a combined model dataset; generating a second model based on the combined model dataset; and determining the accuracy of the second model.
19 . The method of claim 18 , wherein generating the second model is on an ongoing basis.
20 . The method of claim 16 , wherein the data is configured to be separated into the one or more clusters based on an unsupervised machine learning technique.Join the waitlist — get patent alerts
Track US2024046159A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.