Transcript-based insertion of secondary video content into primary video content
Abstract
Certain embodiments involve transcript-based techniques for facilitating insertion of secondary video content into primary video content. For instance, a video editor presents a video editing interface having a primary video section displaying a primary video, a text-based navigation section having navigable portions of a primary video transcript, and a secondary video menu section displaying candidate secondary videos. In some embodiments, candidate secondary videos are obtained by using target terms detected in the transcript to query a remote data source for the candidate secondary videos. In embodiments involving video insertion, the video editor identifies a portion of the primary video corresponding to a portion of the transcript selected within the text-based navigation section. The video editor inserts a secondary video, which is selected from the candidate secondary videos based on an input received at the secondary video menu section, at the identified portion of the primary video.
Claims
exact text as granted — not AI-modified1 . A system comprising:
one or more data storage devices storing a primary video and candidate secondary videos; and a means for inserting secondary video content into primary video content based on selections via a navigable transcript of the primary video content.
2 . The system of claim 1 , further comprising a processing device configured for:
identifying a first time stamp, a second time stamp, and a third time stamp that are associated with playback of the primary video content, wherein the second time stamp corresponds to a portion of the primary video content identified via a text-based navigation section having selectable portions of a transcript; and performing a playback operation that comprises:
retrieving first frames and second frames from a primary video file that includes the primary video content and that lacks any content from the secondary video content;
rendering the first frames from the primary video content for display between the first time stamp of the primary video content and the second time stamp of the primary video content,
rendering frames retrieved from the secondary video content for display starting at the second time stamp and continuing for a duration between the second time stamp and the third time stamp, and
rendering the second frames from the primary video content for display starting from the third time stamp of the primary video content.
3 . The system of claim 3 , further comprising a means for generating a query for candidate secondary videos to be inserted into primary video content based on selections via a navigable transcript of the primary video content.
4 . The system of claim 3 , further comprising a processing device configured for:
accessing a set of words included in a transcript of the primary video content; computing a set of target term probabilities for the set of words, wherein, for each word, a respective target term probability is computed by performing operations comprising:
generating a frequency feature vector representing a frequency of the word within the transcript, a sentiment feature vector representing sentiments associated with the word within the transcript, and a part-of-speech feature vector representing syntaxes of the word within the transcript,
combining the frequency feature vector, the sentiment feature vector, and the part-of-speech feature vector into a target feature vector for the word, and
computing the respective target term probability by applying a recommendation machine-learning model to the target feature vector, wherein the recommendation machine-learning model is trained to associate training target feature vectors with training words tagged as secondary video content search terms in training transcripts; and
selecting, from the set of words, target terms having respective target term probabilities that exceed a threshold probability, wherein one or more of the target terms are included in the query for candidate secondary videos.
5 . The system of claim 1 , further comprising a processing device configured for:
receiving a target term via a search field; generating a candidate video query having a query parameter that includes or is derived from the received target term; retrieving candidate secondary video content by submitting the candidate video query to one or more data sources having the candidate secondary video content; and displaying selectable visual representations of the retrieved candidate secondary video content.
6 . The system of claim 1 , wherein, subsequent to insertion of the selected secondary video content into the primary video content, audio content associated with a portion of the primary video content that is (i) identified via a text-based navigation section having selectable portions of a transcript a portion of the primary video content and (ii) playable with the selected secondary video content is included with the inserted secondary video content.
7 . A non-transitory computer-readable medium storing program code that, when executed by one or more processing devices, causes the one or more processing to perform operations comprising:
inserting secondary video content into primary video content based on selections via a navigable transcript of the primary video content.
8 . The non-transitory computer-readable medium of claim 7 , inserting the secondary video content into the primary video content comprises:
identifying a first time stamp, a second time stamp, and a third time stamp that are associated with playback of the primary video content, wherein the second time stamp corresponds to a portion of the primary video content identified via a text-based navigation section having selectable portions of a transcript; and performing a playback operation that comprises:
retrieving first frames and second frames from a primary video file that includes the primary video content and that lacks any content from the secondary video content;
rendering the first frames from the primary video content for display between the first time stamp of the primary video content and the second time stamp of the primary video content,
rendering frames retrieved from the secondary video content for display starting at the second time stamp and continuing for a duration between the second time stamp and the third time stamp, and
rendering the second frames from the primary video content for display starting from the third time stamp of the primary video content.
9 . The non-transitory computer-readable medium of claim 7 , the operations further comprising generating a query for candidate secondary videos to be inserted into primary video content based on selections via a navigable transcript of the primary video content.
10 . The non-transitory computer-readable medium of claim 9 , the operations further comprising:
accessing a set of words included in a transcript of the primary video content; computing a set of target term probabilities for the set of words, wherein, for each word, a respective target term probability is computed by performing operations comprising:
generating a frequency feature vector representing a frequency of the word within the transcript, a sentiment feature vector representing sentiments associated with the word within the transcript, and a part-of-speech feature vector representing syntaxes of the word within the transcript,
combining the frequency feature vector, the sentiment feature vector, and the part-of-speech feature vector into a target feature vector for the word, and
computing the respective target term probability by applying a recommendation machine-learning model to the target feature vector, wherein the recommendation machine-learning model is trained to associate training target feature vectors with training words tagged as secondary video content search terms in training transcripts; and
selecting, from the set of words, target terms having respective target term probabilities that exceed a threshold probability, wherein one or more of the target terms are included in the query for candidate secondary videos
11 . The non-transitory computer-readable medium of claim 7 , the operations further comprising:
receiving a target term via a search field; generating a candidate video query having a query parameter that includes or is derived from the received target term; retrieving candidate secondary video content by submitting the candidate video query to one or more data sources having the candidate secondary video content; and displaying selectable visual representations of the retrieved candidate secondary video content
12 . The non-transitory computer-readable medium of claim 7 , the operations further comprising, subsequent to insertion of the selected secondary video content into the primary video content, including, with the secondary video content inserted into the primary video content, audio content associated with a portion of the primary video content that is (i) identified via a text-based navigation section having selectable portions of a transcript a portion of the primary video content and (ii) playable with the selected secondary video content.
13 . A method in which one or more processing devices performs operations comprising:
inserting secondary video content into primary video content based on selections via a navigable transcript of the primary video content.
14 . The method of claim 13 , wherein inserting the secondary video into the primary video content comprises:
identifying a first time stamp, a second time stamp, and a third time stamp that are associated with playback of the primary video content, wherein the second time stamp corresponds to a portion of the primary video content identified via a text-based navigation section having selectable portions of a transcript; and performing a playback operation that comprises:
retrieving first frames and second frames from a primary video file that includes the primary video content and that lacks any content from the secondary video content;
rendering the first frames from the primary video content for display between the first time stamp of the primary video content and the second time stamp of the primary video content,
rendering frames retrieved from the secondary video content for display starting at the second time stamp and continuing for a duration between the second time stamp and the third time stamp, and
rendering the second frames from the primary video content for display starting from the third time stamp of the primary video content.
15 . The method of claim 13 , the operations further comprising generating a query for candidate secondary videos to be inserted into primary video content based on selections via a navigable transcript of the primary video content.
16 . The method of claim 15 , the operations further comprising:
accessing a set of words included in a transcript of the primary video content; computing a set of target term probabilities for the set of words, wherein, for each word, a respective target term probability is computed by performing operations comprising:
generating a frequency feature vector representing a frequency of the word within the transcript, a sentiment feature vector representing sentiments associated with the word within the transcript, and a part-of-speech feature vector representing syntaxes of the word within the transcript,
combining the frequency feature vector, the sentiment feature vector, and the part-of-speech feature vector into a target feature vector for the word, and
computing the respective target term probability by applying a recommendation machine-learning model to the target feature vector, wherein the recommendation machine-learning model is trained to associate training target feature vectors with training words tagged as secondary video content search terms in training transcripts; and
selecting, from the set of words, target terms having respective target term probabilities that exceed a threshold probability, wherein one or more of the target terms are included in the query for candidate secondary videos.
17 . The method of claim 13 , the operations further comprising:
receiving a target term via a search field; generating a candidate video query having a query parameter that includes or is derived from the received target term; retrieving candidate secondary video content by submitting the candidate video query to one or more data sources having the candidate secondary video content; and displaying selectable visual representations of the retrieved candidate secondary video content
18 . The method of claim 13 , the operations further comprising, subsequent to insertion of the selected secondary video content into the primary video content, including, with the secondary video content inserted into the primary video content, audio content associated with a portion of the primary video content that is (i) identified via a text-based navigation section having selectable portions of a transcript a portion of the primary video content and (ii) playable with the selected secondary video content.
19 . The method of claim 13 , the operations further comprising:
selecting a portion of the transcript corresponding to a text-selection input received at the text-based navigation section; identifying a portion of the primary video corresponding to the selected portion of the transcript; selecting, from the candidate secondary videos, the secondary video corresponding to a video-selection input received at the secondary video menu section; and
inserting the selected secondary video into the primary video at the identified portion of the primary video.
20 . The method of claim 19 , wherein inserting the selected secondary video into the primary video at the identified portion of the primary video comprises:
identifying a first time stamp, a second time stamp, and a third time stamp that are associated with playback of the primary video, wherein the second time stamp corresponds to the identified portion of the primary video; and performing a playback operation in the video editing interface, wherein the playback operation comprises:
rendering first frames from the primary video for display between the first time stamp of the primary video and the second time stamp of the primary video,
determining that the selected secondary video has been selected for insertion into the primary video,
rendering frames retrieved from the secondary video for display starting at the second time stamp and continuing for a duration between the second time stamp and the third time stamp, and
rendering second frames from the primary video for display starting from the third time stamp of the primary video.Join the waitlist — get patent alerts
Track US2021304799A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.