Systems and Methods for Improved Digital Transcript Creation Using Automated Speech Recognition
Abstract
This disclosure relates generally to systems, methods, and computer readable media for providing improved insights and annotations to enhance recorded audio, video, and/or written transcriptions of testimony. For example, in some embodiments, a method is disclosed for correlating non-verbal cues recognized from an audio and/or video recording of testimony to the corresponding testimony transcript locations. In other embodiments, a method is disclosed for providing testimony-specific artificial intelligence-based insights and annotations to a testimony transcript, e.g., based on the use of machine learning, natural language processing, and/or other techniques. In still other embodiments, a method is disclosed for providing smart citations to a testimony transcript, e.g., which track the location of semantic constructs within the transcript over the course of various modifications being made to the transcript. In yet other embodiments, a method is disclosed for providing intelligent speaker identification-related insights and annotations to an audio recording of a testimony transcript.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
obtaining a first audio recording of a first testimony, wherein the first audio recording comprises one or more first speaking parties; tagging the obtained first audio recording with one or more unique speaker identifiers, wherein each of the one or more unique speaker identifier corresponds to one of the one or more first speaking parties; determining at least one characteristic of the tagged obtained first audio recording; storing the tagged obtained first audio recording and corresponding at least one determined characteristic in a repository; obtaining a second audio recording of a second testimony, wherein the second audio recording comprises one or more second speaking parties; comparing the obtained second audio recording to one or more audio recordings stored in the repository; and in response to finding at least one matching audio recording in the repository, updating the obtained second audio recording with one or more speaker cues, wherein the one or more speaker cues are based, at least in part, on the at least one matching audio recording in the repository.
2 . The method of claim 1 , wherein the audio recording further comprises an audiovisual recording.
3 . The method of claim 1 , wherein one of the at least one matching audio recordings comprises the first audio recording.
4 . The method of claim 1 , wherein one of the one or more speaker cues comprises at least one of the following: a speaker probability value for a voice in the at least one matching audio recording; a unique speaker identifier for a voice in the at least one matching audio recording; a speaker volume indication for a voice in the at least one matching audio recording; and an indication that multiple speakers are likely overlapping each other during a first time interval in the at least one matching audio recording.
5 . The method of claim 1 , wherein one of the at least one characteristics of the tagged obtained first audio recording comprises at least one of the following: a meaning of the obtained first audio recording, an intent of the obtained first audio recording, a source of the obtained first audio recording, and a content of the obtained first audio recording.
6 . The method of claim 1 , wherein there is at least one speaking party in common between the one or more first speaking parties and the one or more second speaking parties.
7 . The method of claim 1 , wherein tagging the obtained first audio recording with one or more unique speaker identifiers comprises tagging the obtained first audio recording based on at least one of the following: input from a human operator; automated voice recognition; automated face recognition; or line level analysis of one or more audio channels of the first audio recording.
8 . A non-transitory program storage device comprising instructions stored thereon to cause one or more processors to:
obtain a first audio recording of a first testimony, wherein the first audio recording comprises one or more first speaking parties; tag the obtained first audio recording with one or more unique speaker identifiers, wherein each of the one or more unique speaker identifier corresponds to one of the one or more first speaking parties; determine at least one characteristic of the tagged obtained first audio recording; store the tagged obtained first audio recording and corresponding at least one determined characteristic in a repository; obtain a second audio recording of a second testimony, wherein the second audio recording comprises one or more second speaking parties; compare the obtained second audio recording to one or more audio recordings stored in the repository; and in response to finding at least one matching audio recording in the repository, update the obtained second audio recording with one or more speaker cues, wherein the one or more speaker cues are based, at least in part, on the at least one matching audio recording in the repository.
9 . The non-transitory program storage device of claim 8 , wherein the audio recording further comprises an audiovisual recording.
10 . The non-transitory program storage device of claim 8 , wherein one of the at least one matching audio recordings comprises the first audio recording.
11 . The non-transitory program storage device of claim 8 , wherein one of the one or more speaker cues comprises at least one of the following: a speaker probability value for a voice in the at least one matching audio recording; a unique speaker identifier for a voice in the at least one matching audio recording; a speaker volume indication for a voice in the at least one matching audio recording; and an indication that multiple speakers are likely overlapping each other during a first time interval in the at least one matching audio recording.
12 . The non-transitory program storage device of claim 8 , wherein one of the at least one characteristics of the tagged obtained first audio recording comprises at least one of the following: a meaning of the obtained first audio recording, an intent of the obtained first audio recording, a source of the obtained first audio recording, and a content of the obtained first audio recording.
13 . The non-transitory program storage device of claim 8 , wherein there is at least one speaking party in common between the one or more first speaking parties and the one or more second speaking parties.
14 . The non-transitory program storage device of claim 8 , wherein the instructions to tag the obtained first audio recording with one or more unique speaker identifiers comprise instructions to tag the obtained first audio recording based on at least one of the following: input from a human operator; automated voice recognition; automated face recognition; or line level analysis of one or more audio channels of the first audio recording.
15 . A device, comprising:
a memory; a display; a user interface; and one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions causing the one or more processors to:
obtain a first audio recording of a first testimony, wherein the first audio recording comprises one or more first speaking parties;
tag the obtained first audio recording with one or more unique speaker identifiers, wherein each of the one or more unique speaker identifier corresponds to one of the one or more first speaking parties;
determine at least one characteristic of the tagged obtained first audio recording;
store the tagged obtained first audio recording and corresponding at least one determined characteristic in a repository;
obtain a second audio recording of a second testimony, wherein the second audio recording comprises one or more second speaking parties;
compare the obtained second audio recording to one or more audio recordings stored in the repository; and
in response to finding at least one matching audio recording in the repository, update the obtained second audio recording with one or more speaker cues, wherein the one or more speaker cues are based, at least in part, on the at least one matching audio recording in the repository.
16 . The device of claim 15 , wherein one of the at least one matching audio recordings comprises the first audio recording.
17 . The device of claim 15 , wherein one of the one or more speaker cues comprises at least one of the following: a speaker probability value for a voice in the at least one matching audio recording; a unique speaker identifier for a voice in the at least one matching audio recording; a speaker volume indication for a voice in the at least one matching audio recording; and an indication that multiple speakers are likely overlapping each other during a first time interval in the at least one matching audio recording.
18 . The device of claim 15 , wherein one of the at least one characteristics of the tagged obtained first audio recording comprises at least one of the following: a meaning of the obtained first audio recording, an intent of the obtained first audio recording, a source of the obtained first audio recording, and a content of the obtained first audio recording.
19 . The device of claim 15 , wherein there is at least one speaking party in common between the one or more first speaking parties and the one or more second speaking parties.
20 . The device of claim 15 , wherein the instructions to tag the obtained first audio recording with one or more unique speaker identifiers comprise instructions to tag the obtained first audio recording based on at least one of the following: input from a human operator; automated voice recognition; automated face recognition; or line level analysis of one or more audio channels of the first audio recording.Join the waitlist — get patent alerts
Track US2020090661A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.