US2020142928A1PendingUtilityA1

Mobile video search

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Oct 21, 2013Filed: Sep 12, 2019Published: May 7, 2020
Est. expiryOct 21, 2033(~7.2 yrs left)· nominal 20-yr term from priority
G06F 16/71G06F 16/7834G06F 16/738G06F 16/7328G06K 9/4671G06K 9/00758G06K 9/00744G06V 10/462G06V 20/46G06V 20/48
58
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A facility for using a mobile device to search video content takes advantage of computing capacity on the mobile device to capture input through a camera and/or a microphone, extract an audio-video signature of the input in real time, and to perform progressive search. By extracting a joint audio-video signature from the input in real time as the input is received and sending the signature to the cloud to search similar video content through the layered audio-video indexing, the facility can provide progressive results of candidate videos for progressive signature captures.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A method comprising:
 receiving, via an input component of a computing device, a plurality of time slices of video content;   extracting audio-video descriptors for the time slices of video content, to obtain aural and visual characteristics of the video content corresponding to the time slice;   generating an audio-video signature associated with one or more of the time slices of video content based at least in part on the audio-video descriptor having been extracted;   providing the audio-video signature associated with the one or more time slices of video content as a query toward a dataset;   receiving candidate results of the query before reaching an end of the time slices of video content before a time window allowed for the query lapses; and   presenting at least some of the candidate results before reaching the end of the time slices of video content.   
     
     
         3 . A method as recited in  claim 2 , wherein the time slices of video content are received from a video output device not associated with the computing device. 
     
     
         4 . A method as recited in  claim 2 , wherein the time slices of video content are received directly or indirectly by at least one of a camera input device or a microphone input device associated with the computing device. 
     
     
         5 . A method as recited in  claim 4 , wherein the time slices of video content are received from a video output device not associated with the computing device. 
     
     
         6 . A method as recited in  claim 2 , wherein a length of individual ones of the plurality of time slices includes at least about 0.1 second and at most about 10.0 seconds. 
     
     
         7 . A method as recited in  claim 2 , wherein the dataset includes a layered audio-video indexed dataset. 
     
     
         8 . A method as recited in  claim 2 , wherein the audio-video signature includes an audio fingerprint and a video hash bit associated with the time slice of video content. 
     
     
         9 . A system configured to perform a method as recited in  claim 2 . 
     
     
         10 . A computer-readable medium having computer-executable instructions encoded thereon, the computer-executable instructions configured to, upon execution, program a device to perform a method as recited in  claim 2 . 
     
     
         11 . A mobile device configured to perform a method as recited in  claim 2 . 
     
     
         12 . A method of layered audio-video search comprising:
 receiving a query audio-video signature related to video content at a layered audio-video engine;   searching a layered audio-video index associated with the layered audio-video engine to identify entries in the layered audio-video index having a similarity to the query audio-video signature above a threshold;   performing geometric verification of respective key frames from the query audio-video signature and entries from the layered audio-video index having the similarity; and   sending candidate results identified via the geometric verification until a window of time allowed for a query using the query audio-video signature has elapsed.   
     
     
         13 . A method as recited in  claim 12 , further comprising progressively processing entries having respective audio-video signatures. 
     
     
         14 . A method as recited in  claim 13 , wherein the progressively processing the entries having respective audio-video signatures includes employing two-part graph-based transformation and matching. 
     
     
         15 . A method as recited in  claim 12 , further comprising:
 determining whether the candidate results are stable; and   determining whether to update the candidate results based at least in part on whether the candidate results are maintained.   
     
     
         16 . A computer-readable medium having computer-executable instructions encoded thereon, the computer-executable instructions configured to, upon execution, program a device to perform operations as recited in  claim 12 . 
     
     
         17 . A system configured to perform a method as recited in  claim 12 . 
     
     
         18 . A computing device configured to perform a method as recited in  claim 12 . 
     
     
         19 . A method of building a layered audio-video index comprising:
 extracting audio-video descriptors corresponding to individual videos in a video dataset until a time window allowed for completing a query using a multi-index has elapsed;   acquiring an audio index, the audio index including audio fingerprints from the audio-video descriptors;   acquiring a visual index, the visual index including visual hash bits from the audio-video descriptors;   creating a first layer including the multi-index by associating the audio index and at least a part of the visual index;   creating a second layer including the visual index; and   maintaining a time relationship between the multi-index of the first layer and the visual index of the second layer.   
     
     
         20 . A method as  claim 9  recites, wherein the at least a part of the visual index for creating the first layer includes a random selection of hash bits from the second layer. 
     
     
         21 . A method as  claim 9  recites, further comprising refining the number of visual points to be searched in the second layer via the audio index.

Join the waitlist — get patent alerts

Track US2020142928A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.