Generating contextually relevant text transcripts of voice recordings within a message thread
Abstract
The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating contextually relevant transcripts of voice recordings based on social networking data. For instance, the disclosed systems receive a voice recording from a user corresponding to a message thread including the user and one or more co-users. The disclosed systems analyze acoustic features of the voice recording to generate transcription-text probabilities. The disclosed systems generate term weights for terms corresponding to objects associated with the user within a social networking system by analyzing user social networking data. Using the contextually aware term weights, the disclosed systems adjust the transcription-text probabilities. Based on the adjusted transcription-text probabilities, the disclosed systems generate a transcript of the voice recording for display within the message thread.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A method comprising:
receiving, from a computing device associated with a user, a voice recording; generating transcription-text probabilities of terms for transcribing the voice recording from acoustic features of the voice recording; determining term weights for the terms from information from nodes linked to the user in a social graph of a social networking system; adjusting the transcription-text probabilities of the terms for transcribing the voice recording based on the term weights; and generating, for display within a message thread, a transcript of the voice recording based on the adjusted transcription-text probabilities of the terms.
22 . The method of claim 21 , further comprising dynamically updating the term weights as changes occur to the social graph of the social networking system or as time passes.
23 . The method of claim 21 , wherein determining the term weights for the terms comprises:
giving a first weight to a first term based on the first term being used in a comment associated with the user in the social graph; and giving a second weight that is greater than the first weight to a second term based on the second term being included within a social networking profile of the user.
24 . The method of claim 21 , wherein determining the term weights for the terms comprises: determining a term weight for a term utilizing a frequency with which the user and one or more co-users mention the term within the message thread.
25 . The method of claim 21 , wherein:
adjusting the transcription-text probabilities of the terms for transcribing the voice recording based on the term weights comprises:
identifying a term or a term unit corresponding to a term weight;
adjusting a transcription-text probability for the term or the term unit from the transcription-text probabilities based on the term weight; and
generating the transcript of the voice recording comprises: selecting the term or the term unit corresponding to the term weight for inclusion in the transcript of the voice recording rather than an alternative term or an alternative term unit with a lower transcription-text probability.
26 . The method of claim 21 , wherein determining the term weights for the terms comprises:
giving a first weight to a first term based on the first term being related to a social networking object with which the user infrequently interacts; and giving a second weight that is greater than the first weight to a second term based on the second term being related to a social networking object with which the user frequently interacts.
27 . The method of claim 21 , wherein determining the term weights for the terms comprises giving more weight to a term that corresponds to a node associated with the user within the social networking system based on:
the term matching a co-user account name linked to the user in the social networking system; the term matching a location linked to the user in the social networking system; the node corresponding to a particular relationship type between the user and one or more co-users in the social networking system; a relatively higher frequency of interaction with the node associated with the term by the user within the social networking system; or a relatively higher frequency of mention by the user of the term within comments of the social networking system.
28 . The method of claim 21 , further comprising:
utilizing an automatic-speech-recognition model to generate the transcription-text probabilities for transcribing the voice recording; and utilizing a trained social-context-machine-learning model to determine the term weights for the terms corresponding to a social networking data associated with the user within the social networking system.
29 . The method of claim 21 , further comprising:
providing, for display within a message user interface corresponding to the message thread, a selectable audio-capture element in relation to a text-input element; and based on detecting a user interaction with the selectable audio-capture element:
capturing the voice recording; and
transmitting the transcript of the voice recording within the message thread to one or more co-users.
30 . A non-transitory computer readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform operations comprising:
receiving, from a computing device associated with a user, a voice recording; generating transcription-text probabilities of terms for transcribing the voice recording from acoustic features of the voice recording; determining term weights for the terms from information from nodes linked to the user in a social graph of a social networking system; adjusting the transcription-text probabilities of the terms for transcribing the voice recording based on the term weights; and generating, for display within a message thread, a transcript of the voice recording based on the adjusted transcription-text probabilities of the terms.
31 . The non-transitory computer readable medium as recited in claim 30 , wherein the operations further comprise dynamically updating the term weights as changes occur to the social graph of the social networking system or as time passes.
32 . The non-transitory computer readable medium as recited in claim 30 , wherein determining the term weights for the terms comprises:
giving a first weight to a first term based on the first term being used in a comment associated with the user in the social graph; and giving a second weight that is greater than the first weight to a second term based on the second term being included within a social networking profile of the user.
33 . The non-transitory computer readable medium as recited in claim 30 , wherein determining the term weights for the terms comprises: determining a term weight for a term utilizing a frequency with which the user and one or more co-users mention the term within the message thread.
34 . The non-transitory computer readable medium as recited in claim 30 , wherein:
adjusting the transcription-text probabilities of the terms for transcribing the voice recording based on the term weights comprises:
identifying a term or a term unit corresponding to a term weight;
adjusting a transcription-text probability for the term or the term unit from the transcription-text probabilities based on the term weight; and
generating the transcript of the voice recording comprises: selecting the term or the term unit corresponding to the term weight for inclusion in the transcript of the voice recording rather than an alternative term or an alternative term unit with a lower transcription-text probability.
35 . A system comprising:
at least one non-transitory computer readable medium comprising a social graph of a social networking system; at least one processor configured to cause the system to:
receive, from a computing device associated with a user, a voice recording;
generate transcription-text probabilities of terms for transcribing the voice recording from acoustic features of the voice recording;
determine term weights for the terms from information from nodes linked to the user in the social graph of the social networking system;
adjust the transcription-text probabilities of the terms for transcribing the voice recording based on the term weights; and
generate, for display within a message thread, a transcript of the voice recording based on the adjusted transcription-text probabilities of the terms.
36 . The system as recited in claim 35 , wherein the at least one processor is further configured to cause the system to generate the term weights by generating priority weights that indicate a likelihood of the user speaking particular terms.
37 . The system as recited in claim 36 , wherein generating the priority weights that indicate the likelihood of the user speaking particular terms is based on a historical usage of the particular terms by the user on the social networking system.
38 . The system as recited in claim 35 , wherein the at least one processor is further configured to cause the system to determine the term weights for the terms by performing operations comprising:
giving a first weight to a first term based on the first term being used in a comment associated with the user in the social graph; and giving a second weight that is greater than the first weight to a second term based on the second term being included within a social networking profile of the user.
39 . The system as recited in claim 35 , wherein the at least one processor is further configured to cause the system to determine the term weights for the terms by performing operations comprising:
giving a first weight to a first term based on the first term being related to a social networking object with which the user infrequently interacts; and giving a second weight that is greater than the first weight to a second term based on the second term being related to a social networking object with which the user frequently interacts.
40 . The system as recited in claim 35 , wherein the at least one processor is further configured to cause the system to:
utilize an automatic-speech-recognition model to generate the transcription-text probabilities for transcribing the voice recording; and utilize a trained social-context-machine-learning model to determine the term weights for the terms corresponding to a social networking data associated with the user within the social networking system.Join the waitlist — get patent alerts
Track US2023223026A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.