Systems and methods for recognizing ambiguity in metadata
Abstract
A method performed at a server system having one or more processors and memory storing one or more programs for execution by the one or more processors includes generating a feature vector that represents a first artist identifier of a plurality of artist identifiers in a first dataset. The feature vector includes a first indication of whether the first artist identifier matches multiple artist entries in one or more second datasets that are distinct from the first dataset. The method includes determining, based at least in part on the first indication, a probability that the first artist identifier is associated with two or more different real-world artists and determining whether the probability satisfies a predetermined probability condition. The method also includes creating, in response to determining that the probability satisfies the predetermined probability condition, a new artist identifier.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
at a server system having one or more processors and memory storing one or more programs for execution by the one or more processors:
generating a feature vector that represents a first artist identifier of a plurality of artist identifiers in a first dataset, the feature vector including a first indication of whether the first artist identifier matches multiple artist entries in one or more second datasets that are distinct from the first dataset;
determining, based at least in part on the first indication, a probability that the first artist identifier is associated with two or more different real-world artists;
determining whether the probability satisfies a predetermined probability condition; and
creating, in response to determining that the probability satisfies the predetermined probability condition, a new artist identifier.
2 . The method of claim 1 , further including:
prior to creating the new artist identifier, consulting the one or more second datasets to identify one or more particular media items that are to be associated with the new artist identifier.
3 . The method of claim 2 , wherein the one or more particular media items were previously associated with the first artist identifier.
4 . The method of claim 2 , wherein identifying the one or more particular media items includes identifying all media items in the one or more second datasets that are associated with the first artist identifier in the first dataset.
5 . The method of claim 4 , wherein identifying all the media items in the one or more second datasets that are associated with the first artist identifier includes:
identifying a first artist entry in the one or more second datasets that is associated with one or more media items and has a same artist name as the first artist identifier; and identifying a second artist entry in the one or more second datasets that is not associated with the one or more media items and has the same artist name as the first artist identifier.
6 . The method of claim 2 , further including:
associating the one or more particular media items with the new artist identifier.
7 . The method of claim 1 , wherein the new artist identifier is associated with a real-world artist in the first dataset.
8 . The method of claim 1 , wherein determining whether the probability satisfies the predetermined probability condition includes determining whether the probability exceeds a predetermined probability threshold.
9 . The method of claim 1 , wherein the first dataset is associated with a media content provider and the one or more second datasets are supplemental databases associated with one or more metadata providers distinct from the media content provider.
10 . The method of claim 1 , wherein the first artist identifier is associated with metadata in the first dataset including a first artist name and one or more media items associated with the first artist name.
11 . The method of claim 10 , wherein the metadata associated with the first artist identifier in the first dataset further includes a title, a country code, a record label, genre, track length, and/or year.
12 . The method of claim 1 , further comprising:
receiving a user input identifying the first artist identifier as ambiguous; and creating the new artist identifier in response to receiving the user input identifying the first artist identifier as ambiguous.
13 . The method of claim 1 , wherein:
the feature vector includes a second indication selected from the group consisting of: (i) an indicator of whether a number of countries of registration of media items associated with the first artist identifier satisfies a country threshold; (ii) an indicator of whether a number of characters in the first artist identifier satisfies a character threshold; (iii) an indicator of whether a number of record labels associated with the first artist identifier satisfies a label threshold; (iv) an indicator of whether the first artist identifier is associated with albums in at least two different languages; and (v) an indicator of whether a difference between an earliest release date and a latest release date of media items associated with the first artist identifier satisfies a time-span threshold; and the probability that the first artist identifier is associated with two or more different real-world artists is determined based at least in part on the first indication and the second indication.
14 . The method of claim 13 , wherein:
the second indication is the indicator of whether the number of countries of registration of media items associated with the first artist identifier satisfies the country threshold; and the method further comprises, at the server system, determining the country threshold based on artist metadata from a training dataset that includes a first set of artist identifiers that are known to be unambiguous and a second set of artist identifiers that are known to be ambiguous.
15 . The method of claim 13 , wherein:
the second indication is the indicator of whether the number of characters in the first artist identifier exceeds the character threshold; and the method further comprises, at the server system, determining the character threshold based on artist metadata from a training dataset that includes a first set of artist identifiers that are known to be unambiguous and a second set of artist identifiers that are known to be ambiguous.
16 . The method of claim 13 , wherein:
the second indication is the indicator of whether the number of record labels associated with the first artist identifier exceeds the label threshold; and the method further comprises, at the server system, determining the label threshold based on artist metadata from a training dataset that includes a first set of artist identifiers that are known to be unambiguous and a second set of artist identifiers that are known to be ambiguous.
17 . The method of claim 1 , wherein determining the probability includes applying a statistical classifier to the feature vector to calculate the probability.
18 . The method of claim 17 , wherein the statistical classifier is a naïve Bayes classifier or a logistic regression classifier.
19 . A server system comprising:
one or more processors; and memory storing one or more programs for execution by the one or more processors, the one or more programs including instructions for:
generating a feature vector that represents a first artist identifier of a plurality of artist identifiers in a first dataset, the feature vector including a first indication of whether the first artist identifier matches multiple artist entries in one or more second datasets that are distinct from the first dataset;
determining, based at least in part on the first indication, a probability that the first artist identifier is associated with two or more different real-world artists;
determining whether the probability satisfies a predetermined probability condition; and
creating, in response to determining that the probability satisfies the predetermined probability condition, a new artist identifier.
20 . A non-transitory computer readable storage medium storing one or more programs, the one or more programs comprising instructions which, when executed by a server system with one or more processors, cause the server system to:
generate a feature vector that represents a first artist identifier of a plurality of artist identifiers in a first dataset, the feature vector including a first indication of whether the first artist identifier matches multiple artist entries in one or more second datasets that are distinct from the first dataset; determine, based at least in part on the first indication, a probability that the first artist identifier is associated with two or more different real-world artists; determine whether the probability satisfies a predetermined probability condition; and create, in response to determining that the probability satisfies the predetermined probability condition, a new artist identifier.Join the waitlist — get patent alerts
Track US2020125981A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.