US2022083583A1PendingUtilityA1

Systems, Methods and Computer Program Products for Associating Media Content Having Different Modalities

Assignee: SPOTIFY ABPriority: Jun 12, 2019Filed: Oct 8, 2021Published: Mar 17, 2022
Est. expiryJun 12, 2039(~12.9 yrs left)· nominal 20-yr term from priority
G06V 20/41G06V 10/806G06V 10/811G06V 10/82G06V 10/454G06F 18/256G06F 18/21355G06F 18/253G06N 3/045G06N 3/0464G06N 3/09G06F 16/65G06F 16/634G06N 3/08G06V 40/20G06F 16/68G06F 16/906G06F 16/438G06F 16/432G06F 16/483G06F 16/908G06K 9/6248G06K 9/00335
59
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems, methods, and computer program products for associating a media content clip(s) with other media content clip(s) having a different modality by determining first embedding vectors of media content items of a first modality, receiving a media content clip of a second modality, determining a second embedding vector of the media content clip of the second modality, ranking the first embedding vectors based on a distance between the embedding vectors and the second embedding vector, and selecting one or more of the media content items of the first modality based on the ranking, thereby pairing media content clips based on emotion.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of associating at least one media content clip with another media content clip having a different modality, the method comprising the steps of:
 determining a plurality of first embedding vectors of a plurality of media content items of a first modality;   receiving a media content clip of a second modality;   determining a second embedding vector of the media content clip of the second modality;   ranking the plurality of first embedding vectors based on a distance between the plurality of first embedding vectors and the second embedding vector; and   selecting one or more of the plurality of media content items of the first modality based on the ranking.   
     
     
         2 . The method according to  claim 1 , wherein the first modality is an auditory modality and the second modality is a visual modality. 
     
     
         3 . The method according to  claim 1 , wherein the first modality is a visual modality and the second modality is an auditory modality. 
     
     
         4 . The method according to  claim 2 , wherein the audio modality is music. 
     
     
         5 . The method according to  claim 3 , wherein the audio modality is music. 
     
     
         6 . The method according to  claim 1 , wherein a model is trained by constraining a stream of video with a plurality of predetermined tags for the first modality and constraining a stream of audio with a plurality of predetermined tags for the second modality. 
     
     
         7 . The method according to  claim 6 , wherein the one or more predetermined tags are used to represent an emotion. 
     
     
         8 . The method according to  claim 7 , wherein the emotions are selected from a set of predetermined emotions. 
     
     
         9 . The method according to  claim 2 , wherein the video modality is obtained from any one of a movie, a television program, a photo, a single frame of a video, or a combination thereof. 
     
     
         10 . The method according to  claim 3 , wherein the video modality is obtained from any one of a movie, a television program, a photo, a single frame of a video, or a combination thereof. 
     
     
         11 . A system configured to associate at least one media content clip with another media content clip having a different modality, the system comprising:
 a computing system including a programmable circuit operatively connected to a memory, the memory storing computer-executable instructions which, when executed by the programmable circuit, cause the computing system to perform:
 determine a plurality of first embedding vectors of a plurality of media content items of a first modality; 
 receive a media content clip of a second modality; 
 determine a second embedding vector of the media content clip of the second modality; 
 rank the plurality of first embedding vectors based on a distance between the plurality of first embedding vectors and the second embedding vector; and 
 select one or more of the plurality of media content items of the first modality based on the ranking. 
   
     
     
         12 . The system according to  claim 11 , wherein the first modality is an auditory modality and the second modality is a visual modality. 
     
     
         13 . The system according to  claim 11 , wherein the first modality is a visual modality and the second modality is an auditory modality. 
     
     
         14 . The system according to  claim 12 , wherein the audio modality is music. 
     
     
         15 . The system according to  claim 13 , wherein the audio modality is music. 
     
     
         16 . The system according to  claim 11 , wherein a model is trained by constraining a stream of video with a plurality of predetermined tags for the first modality and constraining a stream of audio with a plurality of predetermined tags for the second modality. 
     
     
         17 . The system according to  claim 16 , wherein the one or more predetermined tags are used to represent an emotion. 
     
     
         18 . The system according to  claim 17 , wherein the emotions are selected from a set of predetermined emotions. 
     
     
         19 . The system according to  claim 12 , wherein the video modality is obtained from any one of a movie, a television program, a photo, a single frame of a video, or a combination thereof. 
     
     
         20 . The system according to  claim 13 , wherein the video modality is obtained from any one of a movie, a television program, a photo, a single frame of a video, or a combination thereof.

Join the waitlist — get patent alerts

Track US2022083583A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.