Interactive augmentation and integration of real-time speech-to-text
Abstract
In non-limiting examples of the present disclosure, systems, methods and devices for integrating speech-to-text transcription in a productivity application are presented. A request to access a real-time speech-to-text transcription of an audio signal that is being received by a second device is sent by a first device. The real-time speech-to-text transcription may be surfaced in a transcription pane of a productivity application on the first device. A request to translate the transcription to a different language may be received. The transcription may be translated in real-time and surfaced in the transcription pane. A selection of a word in the surfaced transcription may be received. A request to drag the word from the transcription pane and drop it in a window in the productivity application outside of the transcription pane may be received. The word may be surfaced in the window in the productivity application outside of the transcription pane.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A computer-implemented method, comprising:
receiving, from a first user equipment of a speaking user, a request to start a real-time transcription instance for a speech; in response to receiving the request, generating a join code for the real-time transcription instance; and initiating the real-time transcription instance, comprising:
receiving the join code from a plurality of audience user equipment, wherein a first audience user equipment of the plurality of audience user equipment is authenticated via a first user account with a productivity application service that is hosting a productivity application for the first audience user equipment,
receiving an audio stream signal from the first user equipment,
generating, using an artificial intelligence system, a textual transcription of the audio stream signal as the audio stream signal is received, and
transmitting the textual transcription to each of the plurality of audience user equipment as it is generated, wherein transmitting the textual transcription to the first audience user equipment comprises integrating the textual transcription into a transcription pane of the productivity application hosted by the productivity application service.
2 . The computer-implemented method of claim 1 , wherein the first user equipment is authenticated via a second user account with the productivity application service, the method further comprising:
identifying one or more documents stored by the productivity application service associated with the speech; and developing a custom corpus based on the one or more documents, wherein the generating the textual transcription of the audio stream signal comprises using the custom corpus in language processing models of the artificial intelligence system to generate the textual transcription.
3 . The computer-implemented method of claim 2 , wherein generating the textual transcription further comprises:
identifying unique words in the textual transcription that were generated based on the custom corpus; and modifying a format of the unique words in the textual transcription to distinguish the unique words from other words in the textual transcription.
4 . The computer-implemented method of claim 1 , further comprising:
identifying translation settings associated with the first user account; and translating the textual transcription as the textual transcription is generated to generate a translated textual transcription, wherein transmitting the textual transcription to the first audience user equipment comprises transmitting the translated textual transcription as it is generated.
5 . The computer-implemented method of claim 1 , further comprising:
receiving a translation request specifying a language from a second audience user equipment of the plurality of audience user equipment; and translating the textual transcription into the language as the textual transcription is generated to generate a translated textual transcription, wherein transmitting the textual transcription to the second audience user equipment comprises transmitting the translated textual transcription as it is generated.
6 . The computer-implemented method of claim 1 , wherein receiving the join code comprises receiving the join code from a second productivity application executing on a second audience user equipment of the plurality of audience user equipment and wherein transmitting the textual transcription to the second audience user equipment comprises transmitting the textual transcription to the second audience user equipment for integration into the second productivity application.
7 . The computer-implemented method of claim 1 , further comprising:
authenticating the join code from each of the plurality of audience user equipment; and authorizing transmission of the textual transcription based on the authenticating.
8 . A system, comprising:
one or more processors; and a memory having stored thereon instructions that, upon execution by the one or more processors, cause the one or more processors to:
receive, from a first user equipment of a speaking user, a request to start a real-time transcription instance for a speech;
in response to receiving the request, generate a join code for the real-time transcription instance; and
initiate the real-time transcription instance, comprising:
receive the join code from a plurality of audience user equipment,
receive an audio stream signal from the first user equipment,
generate, using an artificial intelligence system, a textual transcription of the audio stream signal as the audio stream signal is received, and
transmit the textual transcription to each of the plurality of audience user equipment as it is generated.
9 . The system of claim 8 , wherein a first audience user equipment of the plurality of audience user equipment is authenticated via a first user account with a productivity application service that is hosting a productivity application for the first audience user equipment, and wherein the instructions to transmit the textual transcription to the first audience user equipment comprises instructions to integrate the textual transcription into a transcription pane of the productivity application hosted by the productivity application service.
10 . The system of claim 9 , wherein the memory comprises further instructions that, upon execution by the one or more processors, cause the one or more processors to:
identify translation settings associated with the first user account; and translate the textual transcription as the textual transcription is generated to generate a translated textual transcription, wherein transmitting the textual transcription to the first audience user equipment comprises transmitting the translated textual transcription as it is generated.
11 . The system of claim 8 , wherein the first user equipment is authenticated via a first user account with a productivity application service, wherein the memory comprises further instructions that, upon execution by the one or more processors, cause the one or more processors to:
identify one or more documents stored by the productivity application service associated with the speech; and develop a custom corpus based on the one or more documents, wherein the instructions to generate the textual transcription of the audio stream signal comprise further instructions that cause the one or more processors to use the custom corpus in language processing models of the artificial intelligence system to generate the textual transcription.
12 . The system of claim 11 , wherein the instructions to generate the textual transcription further comprise instructions that, upon execution by the one or more processors, cause the one or more processors to:
identify unique words in the textual transcription that were generated based on the custom corpus; and modify a format of the unique words in the textual transcription to distinguish the unique words from other words in the textual transcription.
13 . The system of claim 8 , wherein the memory comprises further instructions that, upon execution by the one or more processors, cause the one or more processors to:
receive a translation request specifying a language from a first audience user equipment of the plurality of audience user equipment; and translate the textual transcription into the language as the textual transcription is generated to generate a translated textual transcription, wherein the instructions to transmit the textual transcription to the first audience user equipment comprises instructions to transmit the translated textual transcription as it is generated.
14 . The system of claim 8 , wherein the instructions to receive the join code comprises instructions that cause the one or more processors to receive the join code from a productivity application executing on a first audience user equipment of the plurality of audience user equipment and wherein the instructions to transmit the textual transcription to the first audience user equipment comprises instructions that cause the one or more processors to transmit the textual transcription to the first audience user equipment for integration into the productivity application.
15 . The system of claim 8 , wherein the memory comprises further instructions that, upon execution by the one or more processors, cause the one or more processors to:
authenticate the join code from each of the plurality of audience user equipment; and authorize transmission of the textual transcription based on the authentication.
16 . A computer-readable memory device having stored thereon instructions that, upon execution by one or more processors, cause the one or more processors to:
receive, from a first user equipment of a speaking user, a request to start a real-time transcription instance for a speech; in response to receiving the request, generate a join code for the real-time transcription instance; and initiate the real-time transcription instance, comprising:
receive the join code from a plurality of audience user equipment,
receive an audio stream signal from the first user equipment,
generate, using an artificial intelligence system, a textual transcription of the audio stream signal as the audio stream signal is received, and
transmit the textual transcription to each of the plurality of audience user equipment as it is generated.
17 . The computer-readable memory device of claim 16 , wherein a first audience user equipment of the plurality of audience user equipment is authenticated via a first user account with a productivity application service that is hosting a productivity application for the first audience user equipment, and wherein the instructions to transmit the textual transcription to the first audience user equipment comprises instructions to integrate the textual transcription into a transcription pane of the productivity application hosted by the productivity application service.
18 . The computer-readable memory device of claim 17 , wherein the computer-readable memory device comprises further instructions that, upon execution by the one or more processors, cause the one or more processors to:
identify translation settings associated with the first user account; and translate the textual transcription as the textual transcription is generated to generate a translated textual transcription, wherein transmitting the textual transcription to the first audience user equipment comprises transmitting the translated textual transcription as it is generated.
19 . The computer-readable memory device of claim 16 , wherein the first user equipment is authenticated via a first user account with a productivity application service, wherein the computer-readable memory device comprises further instructions that, upon execution by the one or more processors, cause the one or more processors to:
identify one or more documents stored by the productivity application service associated with the speech; and develop a custom corpus based on the one or more documents, wherein the instructions to generate the textual transcription of the audio stream signal comprise further instructions that cause the one or more processors to use the custom corpus in language processing models of the artificial intelligence system to generate the textual transcription.
20 . The computer-readable memory device of claim 19 , wherein the instructions to generate the textual transcription further comprise instructions that, upon execution by the one or more processors, cause the one or more processors to:
identify unique words in the textual transcription that were generated based on the custom corpus; and modify a format of the unique words in the textual transcription to distinguish the unique words from other words in the textual transcription.Join the waitlist — get patent alerts
Track US2022375463A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.