US2020090661A1PendingUtilityA1

Systems and Methods for Improved Digital Transcript Creation Using Automated Speech Recognition

Assignee: MAGNA LEGAL SERVICES LLCPriority: Sep 13, 2018Filed: Sep 13, 2019Published: Mar 19, 2020
Est. expirySep 13, 2038(~12.1 yrs left)· nominal 20-yr term from priority
G06F 40/35G06F 17/18G10L 15/265G06K 9/00288G06V 40/172G10L 15/26G10L 25/63G10L 17/26G10L 17/10G10L 17/06G10L 17/02
44
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

This disclosure relates generally to systems, methods, and computer readable media for providing improved insights and annotations to enhance recorded audio, video, and/or written transcriptions of testimony. For example, in some embodiments, a method is disclosed for correlating non-verbal cues recognized from an audio and/or video recording of testimony to the corresponding testimony transcript locations. In other embodiments, a method is disclosed for providing testimony-specific artificial intelligence-based insights and annotations to a testimony transcript, e.g., based on the use of machine learning, natural language processing, and/or other techniques. In still other embodiments, a method is disclosed for providing smart citations to a testimony transcript, e.g., which track the location of semantic constructs within the transcript over the course of various modifications being made to the transcript. In yet other embodiments, a method is disclosed for providing intelligent speaker identification-related insights and annotations to an audio recording of a testimony transcript.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method, comprising:
 obtaining a first audio recording of a first testimony, wherein the first audio recording comprises one or more first speaking parties;   tagging the obtained first audio recording with one or more unique speaker identifiers, wherein each of the one or more unique speaker identifier corresponds to one of the one or more first speaking parties;   determining at least one characteristic of the tagged obtained first audio recording;   storing the tagged obtained first audio recording and corresponding at least one determined characteristic in a repository;   obtaining a second audio recording of a second testimony, wherein the second audio recording comprises one or more second speaking parties;   comparing the obtained second audio recording to one or more audio recordings stored in the repository; and   in response to finding at least one matching audio recording in the repository, updating the obtained second audio recording with one or more speaker cues, wherein the one or more speaker cues are based, at least in part, on the at least one matching audio recording in the repository.   
     
     
         2 . The method of  claim 1 , wherein the audio recording further comprises an audiovisual recording. 
     
     
         3 . The method of  claim 1 , wherein one of the at least one matching audio recordings comprises the first audio recording. 
     
     
         4 . The method of  claim 1 , wherein one of the one or more speaker cues comprises at least one of the following: a speaker probability value for a voice in the at least one matching audio recording; a unique speaker identifier for a voice in the at least one matching audio recording; a speaker volume indication for a voice in the at least one matching audio recording; and an indication that multiple speakers are likely overlapping each other during a first time interval in the at least one matching audio recording. 
     
     
         5 . The method of  claim 1 , wherein one of the at least one characteristics of the tagged obtained first audio recording comprises at least one of the following: a meaning of the obtained first audio recording, an intent of the obtained first audio recording, a source of the obtained first audio recording, and a content of the obtained first audio recording. 
     
     
         6 . The method of  claim 1 , wherein there is at least one speaking party in common between the one or more first speaking parties and the one or more second speaking parties. 
     
     
         7 . The method of  claim 1 , wherein tagging the obtained first audio recording with one or more unique speaker identifiers comprises tagging the obtained first audio recording based on at least one of the following: input from a human operator; automated voice recognition; automated face recognition; or line level analysis of one or more audio channels of the first audio recording. 
     
     
         8 . A non-transitory program storage device comprising instructions stored thereon to cause one or more processors to:
 obtain a first audio recording of a first testimony, wherein the first audio recording comprises one or more first speaking parties;   tag the obtained first audio recording with one or more unique speaker identifiers, wherein each of the one or more unique speaker identifier corresponds to one of the one or more first speaking parties;   determine at least one characteristic of the tagged obtained first audio recording;   store the tagged obtained first audio recording and corresponding at least one determined characteristic in a repository;   obtain a second audio recording of a second testimony, wherein the second audio recording comprises one or more second speaking parties;   compare the obtained second audio recording to one or more audio recordings stored in the repository; and   in response to finding at least one matching audio recording in the repository, update the obtained second audio recording with one or more speaker cues, wherein the one or more speaker cues are based, at least in part, on the at least one matching audio recording in the repository.   
     
     
         9 . The non-transitory program storage device of  claim 8 , wherein the audio recording further comprises an audiovisual recording. 
     
     
         10 . The non-transitory program storage device of  claim 8 , wherein one of the at least one matching audio recordings comprises the first audio recording. 
     
     
         11 . The non-transitory program storage device of  claim 8 , wherein one of the one or more speaker cues comprises at least one of the following: a speaker probability value for a voice in the at least one matching audio recording; a unique speaker identifier for a voice in the at least one matching audio recording; a speaker volume indication for a voice in the at least one matching audio recording; and an indication that multiple speakers are likely overlapping each other during a first time interval in the at least one matching audio recording. 
     
     
         12 . The non-transitory program storage device of  claim 8 , wherein one of the at least one characteristics of the tagged obtained first audio recording comprises at least one of the following: a meaning of the obtained first audio recording, an intent of the obtained first audio recording, a source of the obtained first audio recording, and a content of the obtained first audio recording. 
     
     
         13 . The non-transitory program storage device of  claim 8 , wherein there is at least one speaking party in common between the one or more first speaking parties and the one or more second speaking parties. 
     
     
         14 . The non-transitory program storage device of  claim 8 , wherein the instructions to tag the obtained first audio recording with one or more unique speaker identifiers comprise instructions to tag the obtained first audio recording based on at least one of the following: input from a human operator; automated voice recognition; automated face recognition; or line level analysis of one or more audio channels of the first audio recording. 
     
     
         15 . A device, comprising:
 a memory;   a display;   a user interface; and   one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions causing the one or more processors to:
 obtain a first audio recording of a first testimony, wherein the first audio recording comprises one or more first speaking parties; 
 tag the obtained first audio recording with one or more unique speaker identifiers, wherein each of the one or more unique speaker identifier corresponds to one of the one or more first speaking parties; 
 determine at least one characteristic of the tagged obtained first audio recording; 
 store the tagged obtained first audio recording and corresponding at least one determined characteristic in a repository; 
 obtain a second audio recording of a second testimony, wherein the second audio recording comprises one or more second speaking parties; 
 compare the obtained second audio recording to one or more audio recordings stored in the repository; and 
 in response to finding at least one matching audio recording in the repository, update the obtained second audio recording with one or more speaker cues, wherein the one or more speaker cues are based, at least in part, on the at least one matching audio recording in the repository. 
   
     
     
         16 . The device of  claim 15 , wherein one of the at least one matching audio recordings comprises the first audio recording. 
     
     
         17 . The device of  claim 15 , wherein one of the one or more speaker cues comprises at least one of the following: a speaker probability value for a voice in the at least one matching audio recording; a unique speaker identifier for a voice in the at least one matching audio recording; a speaker volume indication for a voice in the at least one matching audio recording; and an indication that multiple speakers are likely overlapping each other during a first time interval in the at least one matching audio recording. 
     
     
         18 . The device of  claim 15 , wherein one of the at least one characteristics of the tagged obtained first audio recording comprises at least one of the following: a meaning of the obtained first audio recording, an intent of the obtained first audio recording, a source of the obtained first audio recording, and a content of the obtained first audio recording. 
     
     
         19 . The device of  claim 15 , wherein there is at least one speaking party in common between the one or more first speaking parties and the one or more second speaking parties. 
     
     
         20 . The device of  claim 15 , wherein the instructions to tag the obtained first audio recording with one or more unique speaker identifiers comprise instructions to tag the obtained first audio recording based on at least one of the following: input from a human operator; automated voice recognition; automated face recognition; or line level analysis of one or more audio channels of the first audio recording.

Join the waitlist — get patent alerts

Track US2020090661A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.