Post-conference playback system having higher perceived quality than originally heard in the conference
Abstract
Some aspects of the present disclosure involve the recording, processing and playback of audio data corresponding to conferences, such as teleconferences. In some teleconference implementations, the audio experience heard when a recording of the conference is played back may be substantially different from the audio experience of an individual conference participant during the original teleconference. In some implementations, the recorded audio data may include at least some audio data that was not available during the teleconference. In some examples, the spatial characteristics of the played-back audio data may be different from that of the audio heard by participants of the teleconference.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . An apparatus, comprising:
an interface system; and a control system configured to perform operations comprising:
receiving, via the interface system, recorded audio data for a teleconference, the recorded audio data including an individual uplink data packet stream corresponding to a telephone endpoint used by one or more teleconference participants;
analyzing sequence number data of data packets in the individual uplink data packet stream, wherein the analyzing involves determining whether the individual uplink data packet stream includes at least one out-of-order data packet; and
re-ordering the individual uplink data packet stream according to the sequence number data if the uplink data packet stream includes at least one out-of-order data packet.
2 . The apparatus of claim 1 , wherein at least one data packet of the individual uplink data packet stream was received after a mouth-to-ear latency time threshold of the teleconference.
3 . The apparatus of claim 1 , wherein the operations comprise:
receiving, via the interface system, teleconference metadata; and indexing the individual uplink data packet stream based, at least in part, on the teleconference metadata.
4 . The apparatus of claim 1 , wherein the recorded audio data includes a plurality of individual encoded uplink data packet streams, each of the individual encoded uplink data packet streams corresponding to a telephone endpoint used by one or more teleconference participants, wherein the control system includes a joint analysis module capable of analyzing a plurality of individual uplink data packet streams and wherein the operations comprise:
decoding the plurality of individual encoded uplink data packet streams; and providing a plurality of individual decoded uplink data packet streams to the joint analysis module.
5 . The apparatus of claim 4 , wherein the control system includes a speech recognition module capable of recognizing speech and generating speech recognition results data, wherein the control system is capable of providing one or more individual decoded uplink data packet streams to the speech recognition module and wherein the speech recognition module is further capable of providing the speech recognition results data to the joint analysis module.
6 . The apparatus of claim 5 , wherein the joint analysis module is configured to perform operations comprising:
identifying keywords in the speech recognition results data; and indexing keyword locations.
7 . The apparatus of claim 4 , wherein the control system includes a speaker diarization module, wherein the control system is capable of providing an individual decoded uplink data packet stream to the speaker diarization module and wherein the speaker diarization module is configured to perform operations comprising:
identifying speech of each of multiple teleconference participants in an individual decoded uplink data packet stream; generating a speaker diary indicating times at which each of the multiple teleconference participants were speaking; and providing the speaker diary to the joint analysis module.
8 . The apparatus of claim 4 , wherein the joint analysis module is further configured to perform operations comprising determining conversational dynamics data, the conversational dynamics data including one or more types of data selected from a group of data types consisting of: data indicating the frequency and duration of conference participant speech; data indicating instances of conference participant doubletalk during which at least two conference participants are speaking simultaneously; and data indicating instances of conference participant conversations.
9 . An apparatus, comprising:
an interface system; and control means for:
receiving, via the interface system, recorded audio data for a teleconference, the recorded audio data including an individual uplink data packet stream corresponding to a telephone endpoint used by one or more teleconference participants;
analyzing sequence number data of data packets in the individual uplink data packet stream, wherein the analyzing involves determining whether the individual uplink data packet stream includes at least one out-of-order data packet; and
re-ordering the individual uplink data packet stream according to the sequence number data if the uplink data packet stream includes at least one out-of-order data packet.
10 . The apparatus of claim 9 , wherein the recorded audio data includes a plurality of individual encoded uplink data packet streams, each of the individual encoded uplink data packet streams corresponding to a telephone endpoint used by one or more teleconference participants, wherein the control means includes a joint analysis module capable of analyzing a plurality of individual uplink data packet streams and wherein the control means includes means for:
decoding the plurality of individual encoded uplink data packet streams; and providing a plurality of individual decoded uplink data packet streams to the joint analysis module.
11 . The apparatus of claim 10 , wherein the control means includes a speech recognition module capable of recognizing speech and generating speech recognition results data, wherein the control means includes means for providing one or more individual decoded uplink data packet streams to the speech recognition module and wherein the speech recognition module is further capable of providing the speech recognition results data to the joint analysis module.
12 . The apparatus of claim 11 , wherein the joint analysis module is capable of:
identifying keywords in the speech recognition results data; and indexing keyword locations.
13 . The apparatus of claim 10 , wherein the control means includes a speaker diarization module, wherein the control means includes means for providing an individual decoded uplink data packet stream to the speaker diarization module and wherein the speaker diarization module is capable of:
identifying speech of each of multiple teleconference participants in an individual decoded uplink data packet stream; generating a speaker diary indicating times at which each of the multiple teleconference participants were speaking; and providing the speaker diary to the joint analysis module.
14 . The apparatus of claim 10 , wherein the joint analysis module is further capable of determining conversational dynamics data, the conversational dynamics data including one or more types of data selected from a group of data types consisting of: data indicating the frequency and duration of conference participant speech; data indicating instances of conference participant doubletalk during which at least two conference participants are speaking simultaneously; and data indicating instances of conference participant conversations.
15 . A method of processing audio data, comprising:
receiving, via an interface system, recorded audio data for a teleconference, the recorded audio data including an individual uplink data packet stream corresponding to a telephone endpoint used by one or more teleconference participants; analyzing sequence number data of data packets in the individual uplink data packet stream, wherein the analyzing involves determining whether the individual uplink data packet stream includes at least one out-of-order data packet; and re-ordering the individual uplink data packet stream according to the sequence number data if the uplink data packet stream includes at least one out-of-order data packet.
16 . The method of claim 15 , wherein the recorded audio data includes a plurality of individual encoded uplink data packet streams, each of the individual encoded uplink data packet streams corresponding to a telephone endpoint used by one or more teleconference participants, further comprising:
decoding the plurality of individual encoded uplink data packet streams; and analyzing the plurality of individual uplink data packet streams.
17 . A non-transitory medium having software stored thereon, the software including instructions for processing audio data by causing at least one device to perform operations:
receiving, via an interface system, recorded audio data for a teleconference, the recorded audio data including an individual uplink data packet stream corresponding to a telephone endpoint used by one or more teleconference participants;
analyzing sequence number data of data packets in the individual uplink data packet stream, wherein the analyzing involves determining whether the individual uplink data packet stream includes at least one out-of-order data packet; and
re-ordering the individual uplink data packet stream according to the sequence number data if the uplink data packet stream includes at least one out-of-order data packet.
18 . The non-transitory medium of claim 17 , wherein the recorded audio data includes a plurality of individual encoded uplink data packet streams, each of the individual encoded uplink data packet streams corresponding to a telephone endpoint used by one or more teleconference participants, wherein the software includes instructions for:
decoding the plurality of individual encoded uplink data packet streams; and analyzing the plurality of individual uplink data packet streams.
19 . The non-transitory medium of claim 18 , wherein the software includes instructions for:
recognizing speech in one or more individual decoded uplink data packet streams; and generating speech recognition results data.
20 . The non-transitory medium of claim 19 , wherein the software includes instructions for:
identifying keywords in the speech recognition results data; and indexing keyword locations.
21 . The non-transitory medium of claim 18 , wherein the software includes instructions for:
identifying speech of each of multiple teleconference participants in an individual decoded uplink data packet stream; and generating a speaker diary indicating times at which each of the multiple teleconference participants were speaking.
22 . The non-transitory medium of claim 18 , wherein analyzing the plurality of individual uplink data packet streams involves determining conversational dynamics data, the conversational dynamics data including one or more types of data selected from a group of data types consisting of: data indicating the frequency and duration of conference participant speech; data indicating instances of conference participant doubletalk during which at least two conference participants are speaking simultaneously; and data indicating instances of conference participant conversations.Join the waitlist — get patent alerts
Track US2020127865A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.